Skip to contents

Generates a two-level VAR(1) process for N persons over n_prompts prompts and then deletes whole prompts according to one or more missingness motifs. Response probabilities use a probit link. Intercepts are calibrated numerically so that the realized response rate matches compliance.

Usage

simulate_ema(
  N = 100,
  n_prompts = 56,
  params = default_params(),
  motifs = "M0",
  compliance = 0.75,
  gamma = -0.5,
  delta = -1,
  kappa_R = 1,
  burden = 0,
  rho_propensity = 0.4,
  sd_propensity = 0.5,
  p_context = 0.3,
  rho_context = 0,
  gamma_C = 0.6,
  kappa_C = 1,
  rho_react = 0.2,
  sensor_cor = NULL,
  p_probe = 0,
  burnin = 50,
  seed = NULL,
  days = NULL
)

Arguments

N

Number of persons.

n_prompts

Number of prompts per person (called T in the archived versions 0.2.x of the package).

params

List with Phi, Psi, mu, Sigma_mu (see default_params).

motifs

Character vector of motif codes (see dm_graph).

compliance

Target overall response rate.

gamma

Coefficient vector of X_{t-1} in the response model (motif M1); a scalar is applied to the first variable.

delta

Coefficient vector of X_t in the response model (motif M2); a scalar is applied to the first variable. Negative values mean that high states are skipped. Both gamma and delta act on the raw (uncentered) state, so a self-censoring person with a high typical state also responds less often overall.

kappa_R

Coefficient of R_{t-1} in the response model (M3); the term enters as kappa_R * (R_{t-1} - compliance), so the propensity after an answered prompt exceeds that after a skipped prompt by kappa_R probit units.

burden

Coefficient of a linear time trend in the response model (M3); the term enters as burden * trend, where the trend runs from -0.5 to 0.5 over the burn-in and the recorded prompts together, so a negative value gives declining compliance over the study and the realized change over the recorded prompts is about burden * n_prompts / (n_prompts + burnin) probit units.

rho_propensity

Correlation between the person-level response propensity and the person's mean on the first variable (M4).

sd_propensity

Standard deviation of the person-level propensity (M4), in probit units.

p_context

Probability that the (binary) context C_t = 1 (M5).

rho_context

Persistence of the context: with this probability C_t copies C_{t-1}, otherwise it is drawn afresh with probability p_context; 0 gives serially independent contexts (the marginal rate is p_context either way).

gamma_C

Shift of the state vector when C_t = 1 (M5); a scalar is applied to the first variable.

kappa_C

Coefficient of C_t in the response model (M5).

rho_react

Shift of the state vector at prompt t when the previous prompt was answered (M6, reactivity); the term enters as rho_react * (R_{t-1} - compliance), so the contrast between an answered and a skipped predecessor is rho_react; a scalar applies to the first variable.

sensor_cor

Correlation between an always-observed sensor channel and the within-person fluctuation of the first state variable (its deviation from the person mean, standardized); NULL for no sensor.

p_probe

Probability that a prompt is a randomized probe that forces a response; 0 for no probes.

burnin

Number of burn-in prompts discarded before recording.

seed

Optional integer seed (the caller's random-number state is restored afterwards).

days

Optional number of prompts per day; when given, a day column is added (the response model itself does not use the day).

Value

A list of class "ema_sim" with data (long data frame:

id, time, optional day, R, the state columns with NA at skipped prompts, and the optional columns C

(context, M5), S (sensor) and Z (probe)), full

(the state columns without deletion, with id and time),

mu_i (the true person means), alpha0 (the calibrated response intercept), a_i (the person propensities, M4),

params, motifs, settings (the design and the response-model arguments) and vars. It has a print method.

See also

dm_graph for the same motifs as a graph, simulate_from_fit to simulate from a fitted model, default_params.

Examples

sim <- simulate_ema(N = 20, n_prompts = 15, motifs = c("M2", "M4"), compliance = 0.7, seed = 1)
sim
#> Simulated EMA data: 20 persons x 15 prompts, 4 variables
#>   motifs: M2 + M4   realized response rate: 0.7 
head(sim$data)
#>   id time R      NegA     PosA    Stress  Fatigue
#> 1  1    1 1 2.4459972 3.786772 0.8344344 4.646153
#> 2  1    2 1 0.6384173 3.827842 0.4050047 3.053483
#> 3  1    3 1 0.9811262 3.991629 1.3967747 3.760117
#> 4  1    4 1 1.5826669 4.652583 1.9895342 2.736375
#> 5  1    5 0        NA       NA        NA       NA
#> 6  1    6 0        NA       NA        NA       NA
# the deleted states are kept for checking estimators
head(sim$full)
#>   id time      NegA     PosA    Stress  Fatigue
#> 1  1    1 2.4459972 3.786772 0.8344344 4.646153
#> 2  1    2 0.6384173 3.827842 0.4050047 3.053483
#> 3  1    3 0.9811262 3.991629 1.3967747 3.760117
#> 4  1    4 1.5826669 4.652583 1.9895342 2.736375
#> 5  1    5 2.7230413 3.811979 2.8471250 3.936757
#> 6  1    6 3.4129725 3.551051 1.2580618 4.358515
# a sensor, probes, a persistent observed context and a day structure
sim2 <- simulate_ema(N = 20, n_prompts = 20, motifs = c("M2", "M5"), sensor_cor = 0.6,
                     p_probe = 0.1, rho_context = 0.5, days = 5, seed = 2)
names(sim2$data)
#>  [1] "id"      "time"    "day"     "R"       "NegA"    "PosA"    "Stress" 
#>  [8] "Fatigue" "C"       "S"       "Z"      
mean(sim2$data$R[sim2$data$Z == 1])   # probes are always answered
#> [1] 1