Simulate experience-sampling data with a declared missingness mechanism
Source:R/simulate.R
simulate_ema.RdGenerates a two-level VAR(1) process for N persons over n_prompts
prompts and then deletes whole prompts according to one or more
missingness motifs. Response probabilities use a probit link. Intercepts
are calibrated numerically so that the realized response rate matches
compliance.
Usage
simulate_ema(
N = 100,
n_prompts = 56,
params = default_params(),
motifs = "M0",
compliance = 0.75,
gamma = -0.5,
delta = -1,
kappa_R = 1,
burden = 0,
rho_propensity = 0.4,
sd_propensity = 0.5,
p_context = 0.3,
rho_context = 0,
gamma_C = 0.6,
kappa_C = 1,
rho_react = 0.2,
sensor_cor = NULL,
p_probe = 0,
burnin = 50,
seed = NULL,
days = NULL
)Arguments
- N
Number of persons.
- n_prompts
Number of prompts per person (called
Tin the archived versions 0.2.x of the package).- params
List with
Phi,Psi,mu,Sigma_mu(seedefault_params).- motifs
Character vector of motif codes (see
dm_graph).- compliance
Target overall response rate.
- gamma
Coefficient vector of
X_{t-1}in the response model (motif M1); a scalar is applied to the first variable.- delta
Coefficient vector of
X_tin the response model (motif M2); a scalar is applied to the first variable. Negative values mean that high states are skipped. Bothgammaanddeltaact on the raw (uncentered) state, so a self-censoring person with a high typical state also responds less often overall.- kappa_R
Coefficient of
R_{t-1}in the response model (M3); the term enters askappa_R * (R_{t-1} - compliance), so the propensity after an answered prompt exceeds that after a skipped prompt bykappa_Rprobit units.- burden
Coefficient of a linear time trend in the response model (M3); the term enters as
burden * trend, where the trend runs from -0.5 to 0.5 over the burn-in and the recorded prompts together, so a negative value gives declining compliance over the study and the realized change over the recorded prompts is aboutburden * n_prompts / (n_prompts + burnin)probit units.- rho_propensity
Correlation between the person-level response propensity and the person's mean on the first variable (M4).
- sd_propensity
Standard deviation of the person-level propensity (M4), in probit units.
- p_context
Probability that the (binary) context
C_t = 1(M5).- rho_context
Persistence of the context: with this probability
C_tcopiesC_{t-1}, otherwise it is drawn afresh with probabilityp_context;0gives serially independent contexts (the marginal rate isp_contexteither way).- gamma_C
Shift of the state vector when
C_t = 1(M5); a scalar is applied to the first variable.- kappa_C
Coefficient of
C_tin the response model (M5).- rho_react
Shift of the state vector at prompt
twhen the previous prompt was answered (M6, reactivity); the term enters asrho_react * (R_{t-1} - compliance), so the contrast between an answered and a skipped predecessor isrho_react; a scalar applies to the first variable.- sensor_cor
Correlation between an always-observed sensor channel and the within-person fluctuation of the first state variable (its deviation from the person mean, standardized);
NULLfor no sensor.- p_probe
Probability that a prompt is a randomized probe that forces a response;
0for no probes.- burnin
Number of burn-in prompts discarded before recording.
- seed
Optional integer seed (the caller's random-number state is restored afterwards).
- days
Optional number of prompts per day; when given, a
daycolumn is added (the response model itself does not use the day).
Value
A list of class "ema_sim" with data (long data frame:
id, time, optional day, R, the state columns
with NA at skipped prompts, and the optional columns C
(context, M5), S (sensor) and Z (probe)), full
(the state columns without deletion, with id and time),
mu_i (the true person means), alpha0 (the calibrated
response intercept), a_i (the person propensities, M4),
params, motifs, settings (the design and the
response-model arguments) and vars. It has a print method.
See also
dm_graph for the same motifs as a graph,
simulate_from_fit to simulate from a fitted model,
default_params.
Examples
sim <- simulate_ema(N = 20, n_prompts = 15, motifs = c("M2", "M4"), compliance = 0.7, seed = 1)
sim
#> Simulated EMA data: 20 persons x 15 prompts, 4 variables
#> motifs: M2 + M4 realized response rate: 0.7
head(sim$data)
#> id time R NegA PosA Stress Fatigue
#> 1 1 1 1 2.4459972 3.786772 0.8344344 4.646153
#> 2 1 2 1 0.6384173 3.827842 0.4050047 3.053483
#> 3 1 3 1 0.9811262 3.991629 1.3967747 3.760117
#> 4 1 4 1 1.5826669 4.652583 1.9895342 2.736375
#> 5 1 5 0 NA NA NA NA
#> 6 1 6 0 NA NA NA NA
# the deleted states are kept for checking estimators
head(sim$full)
#> id time NegA PosA Stress Fatigue
#> 1 1 1 2.4459972 3.786772 0.8344344 4.646153
#> 2 1 2 0.6384173 3.827842 0.4050047 3.053483
#> 3 1 3 0.9811262 3.991629 1.3967747 3.760117
#> 4 1 4 1.5826669 4.652583 1.9895342 2.736375
#> 5 1 5 2.7230413 3.811979 2.8471250 3.936757
#> 6 1 6 3.4129725 3.551051 1.2580618 4.358515
# a sensor, probes, a persistent observed context and a day structure
sim2 <- simulate_ema(N = 20, n_prompts = 20, motifs = c("M2", "M5"), sensor_cor = 0.6,
p_probe = 0.1, rho_context = 0.5, days = 5, seed = 2)
names(sim2$data)
#> [1] "id" "time" "day" "R" "NegA" "PosA" "Stress"
#> [8] "Fatigue" "C" "S" "Z"
mean(sim2$data$R[sim2$data$Z == 1]) # probes are always answered
#> [1] 1