Full-information maximum likelihood under missing at random (state-space EM)
Source:R/fit_fiml.R
fit_fiml.RdFits the two-level VAR(1) with random person means by maximum likelihood, treating skipped prompts as missing at random. The likelihood is computed by the Kalman filter over the augmented state \((X_t, \mu_i)\) and maximized by the EM algorithm of Shumway and Stoffer (1982). This is the likelihood that dynamic structural equation modeling and related software maximize (with a diffuse or informative prior instead of the flat likelihood), and it is the "default pipeline" against which Yu (2026) measures the damage of state-dependent nonresponse.
Usage
fit_fiml(
data,
vars,
id = "id",
time = "time",
R = "R",
start = NULL,
max_iter = 200L,
tol = 1e-05,
verbose = FALSE
)Arguments
- data
Long data frame with one row per scheduled prompt.
- vars
Names of the state columns.
- id, time
Names of the person and prompt-index columns.
- R
Name of the response-indicator column (0/1, no
NA). If the default name is not among the columns ofdata, a prompt counts as answered when allvarsare non-missing; any other name must exist.- start
Optional list with
Phi,Psi,mu,Sigma_mu; by default the answered-pairs estimates are used.- max_iter, tol
EM control.
- verbose
Print the EM trace.
Value
An object of class "fiml_fit": a list with Phi,
Psi, pcor, mu, Sigma_mu, Sigma0
(the stationary covariance of the within-person deviations), loglik
(the trace of log-likelihood values over the EM iterations),
iterations, converged, N, n_prompts (the
number of prompts spanned by the prompt index) and vars. It has a print method.
Details
The prompts of a person are treated as one equally spaced series
(the overnight transition is a lag-1 transition, as in the default
specification of dynamic structural equation models with equally spaced
occasions); there is no day argument. Standard errors are not
computed (the estimator serves as the MAR benchmark in the simulation
studies of Yu, 2026). loglik holds the log-likelihood evaluated at the
parameters entering each EM iteration, so its last element belongs to the
penultimate iterate; loglik_fiml evaluates the returned
estimates.
References
Shumway, R. H., & Stoffer, D. S. (1982). An approach to time series smoothing and forecasting using the EM algorithm. Journal of Time Series Analysis, 3, 253-264. doi:10.1111/j.1467-9892.1982.tb00349.x
Examples
sim <- simulate_ema(N = 30, n_prompts = 20, motifs = "M1", seed = 1)
f <- fit_fiml(sim$data, sim$vars)
f
#> FIML (MAR) two-level VAR(1) by state-space EM: 30 persons, 20 prompts; 16 iterations, converged
#> Lagged coefficients Phi:
#> NegA PosA Stress Fatigue
#> NegA 0.249 0.035 0.255 -0.065
#> PosA -0.083 0.362 -0.027 -0.002
#> Stress 0.104 -0.011 0.501 0.047
#> Fatigue -0.093 -0.196 -0.046 0.328
#> Contemporaneous partial correlations:
#> NegA PosA Stress Fatigue
#> NegA 1.000 -0.349 0.314 -0.041
#> PosA -0.349 1.000 0.066 -0.087
#> Stress 0.314 0.066 1.000 0.273
#> Fatigue -0.041 -0.087 0.273 1.000
#> Means: 2.429 4.059 2.841 3.012
f$converged; length(f$loglik)
#> [1] TRUE
#> [1] 16
# under self-censoring the MAR likelihood is biased like the answered-pairs estimator
sim2 <- simulate_ema(N = 30, n_prompts = 20, motifs = "M2", delta = -1, seed = 1)
c(fiml = fit_fiml(sim2$data, sim2$vars)$Phi[1, 1],
pairs = fit_pairs(sim2$data, sim2$vars)$Phi[1, 1],
truth = default_params()$Phi[1, 1])
#> fiml pairs truth
#> 0.2375767 0.1977559 0.4000000