Skip to contents

Estimates the lagged-coefficient matrix \(\Phi\), the innovation covariance \(\Psi\), person intercepts and person means from complete adjacent prompt pairs, with person-specific intercepts (the within estimator). This is the estimator that recovers the transition kernel under the recoverable motifs (M0, M1, M3, M4 and their combinations, and M5 with an observed context; see recoverability). Optional weights turn it into the inverse-probability-weighted or tilted estimator. By default the half-panel (split-panel) jackknife of Dhaene and Jochmans (2015) removes the finite-\(T\) (Nickell) bias of the within estimator of \(\Phi\).

Usage

fit_pairs(
  data,
  vars,
  id = "id",
  time = "time",
  day = NULL,
  R = "R",
  weights = NULL,
  pairs = NULL,
  min_pairs = 5L,
  bias_correct = TRUE,
  covariates = NULL
)

Arguments

data

Long data frame with one row per scheduled prompt.

vars

Names of the state columns.

id, time

Names of the person and prompt-index columns.

day

Optional name of a day column; pairs are formed only within a day (the overnight gap is not a lag-1 transition).

R

Name of the response-indicator column (0/1, no NA). If the default name is not among the columns of data, a prompt counts as answered when all vars are non-missing; any other name must exist.

weights

Optional non-negative weights, one per complete pair (in the order returned by make_pairs).

pairs

Optional precomputed result of make_pairs.

min_pairs

Persons with fewer complete pairs are dropped from the between-person summaries (they still contribute to \(\Phi\)). A person can have at most \(T - 1\) pairs, so min_pairs must be smaller than the number of prompts per person.

bias_correct

Logical; apply the half-panel jackknife to \(\Phi\), \(\Psi\), the person intercepts and means, and the between-person covariance (all of which carry \(O(1/T)\) incidental-parameter bias).

covariates

Optional names of observed context columns measured at prompt \(t\) that enter the within regression as covariates; the returned \(\Phi\) is then the context-conditional transition kernel (the recoverable estimand under M5 with an observed context), and B_cov holds the covariate coefficients. Person means are evaluated at each person's mean covariate level.

Value

An object of class "pairs_fit": a list with Phi

(lagged coefficients, rows = outcome at \(t\), columns = predictor at

\(t-1\)), Psi (innovation covariance), pcor

(contemporaneous partial correlations), B_cov (covariate coefficients, or NULL), intercepts (person intercepts

\(c_i\)), mu_dyn (dynamics-recovered person means

\((I - \Phi)^{-1} c_i\)), mu_obs (observed person means),

mu and Sigma_mu (person-weighted between-person law based on mu_dyn), mu_obs_person (person-weighted mean of the observed means), mu_obs_pooled (prompt-weighted mean of answered prompts), Sigma_mu_obs, the uncorrected (least-squares) versions

Phi_ls, Psi_ls, mu_ls, Sigma_mu_ls,

vcov_Phi (cluster-robust covariance of the stacked rows of

Phi), n_pairs, n_persons (persons with at least one complete pair), n_pairs_person,

n_dropped, weights, pairs (the

make_pairs object), vars, residuals and

bias_correct. Methods: print, summary.pairs_fit,

coef, vcov, confint, nobs.

Details

Person intercepts are handled by weighted within-person demeaning on the sample of complete pairs (equivalent to person dummies), not by centering on the observed person means as in the two-step mlVAR approach; the two coincide for \(\Phi\) only in large samples. Persons contribute to \(\Phi\) and \(\Psi\) whatever their number of pairs; the between-person summaries use persons with at least min_pairs pairs, and the number dropped is reported in n_dropped.

References

Dhaene, G., & Jochmans, K. (2015). Split-panel jackknife estimation of fixed-effect models. The Review of Economic Studies, 82, 991-1030. doi:10.1093/restud/rdv007

Nickell, S. (1981). Biases in dynamic models with fixed effects. Econometrica, 49, 1417-1426. doi:10.2307/1911408

Examples

sim <- simulate_ema(N = 40, n_prompts = 30, motifs = "M1", seed = 1)
f <- fit_pairs(sim$data, sim$vars)
f
#> Within-person VAR(1) from 674 complete adjacent pairs, 40 persons (half-panel jackknife) 
#> Lagged coefficients Phi (rows = outcome at t, columns = predictor at t-1):
#>           NegA   PosA Stress Fatigue
#> NegA     0.377 -0.020  0.186  -0.014
#> PosA    -0.057  0.441 -0.069   0.104
#> Stress   0.117 -0.026  0.436   0.143
#> Fatigue -0.001 -0.097  0.026   0.341
#> Contemporaneous partial correlations (innovations):
#>           NegA   PosA Stress Fatigue
#> NegA     1.000 -0.392  0.325  -0.007
#> PosA    -0.392  1.000  0.065  -0.043
#> Stress   0.325  0.065  1.000   0.182
#> Fatigue -0.007 -0.043  0.182   1.000
#> Between-person means (dynamics-recovered): 2.51 3.839 2.902 2.976 
summary(f)
#>                coef     estimate         se        lower       upper
#> 1        NegA<-NegA  0.377209619 0.04546763  0.285242657  0.46917658
#> 2        NegA<-PosA -0.020148335 0.03810703 -0.097227071  0.05693040
#> 3      NegA<-Stress  0.185641404 0.04502284  0.094574114  0.27670869
#> 4     NegA<-Fatigue -0.014290402 0.04203572 -0.099315662  0.07073486
#> 5        PosA<-NegA -0.057096878 0.05102948 -0.160313753  0.04612000
#> 6        PosA<-PosA  0.440651793 0.03898408  0.361799058  0.51950453
#> 7      PosA<-Stress -0.068794998 0.06108369 -0.192348420  0.05475842
#> 8     PosA<-Fatigue  0.104248748 0.05096079  0.001170818  0.20732668
#> 9      Stress<-NegA  0.117282591 0.04846479  0.019253305  0.21531188
#> 10     Stress<-PosA -0.025802913 0.03436702 -0.095316772  0.04371095
#> 11   Stress<-Stress  0.435789576 0.04305469  0.348703243  0.52287591
#> 12  Stress<-Fatigue  0.142998521 0.02936493  0.083602335  0.20239471
#> 13    Fatigue<-NegA -0.001116763 0.05318874 -0.108701153  0.10646763
#> 14    Fatigue<-PosA -0.096759728 0.03937903 -0.176411343 -0.01710811
#> 15  Fatigue<-Stress  0.026294000 0.03793284 -0.050432415  0.10302041
#> 16 Fatigue<-Fatigue  0.340532511 0.04538356  0.248735593  0.43232943
#>              z            p
#> 1   8.29622348 3.826823e-10
#> 2  -0.52873017 5.999890e-01
#> 3   4.12327175 1.890931e-04
#> 4  -0.33995857 7.357119e-01
#> 5  -1.11889978 2.700273e-01
#> 6  11.30337941 7.127472e-14
#> 7  -1.12624172 2.669459e-01
#> 8   2.04566579 4.757589e-02
#> 9   2.41995471 2.028256e-02
#> 10 -0.75080449 4.572768e-01
#> 11 10.12176751 1.814320e-12
#> 12  4.86970342 1.887754e-05
#> 13 -0.02099623 9.833557e-01
#> 14 -2.45713818 1.855680e-02
#> 15  0.69317242 4.923090e-01
#> 16  7.50343291 4.407210e-09
coef(f)[1:4]; confint(f)[1:2, ]
#>    NegA<-NegA    NegA<-PosA  NegA<-Stress NegA<-Fatigue 
#>    0.37720962   -0.02014833    0.18564140   -0.01429040 
#>                  2.5 %    97.5 %
#> NegA<-NegA  0.28524266 0.4691766
#> NegA<-PosA -0.09722707 0.0569304
# compare with the population values used by the simulator
round(f$Phi - default_params()$Phi, 2)
#>          NegA  PosA Stress Fatigue
#> NegA    -0.02 -0.02   0.04   -0.01
#> PosA     0.04  0.09  -0.07    0.10
#> Stress  -0.03 -0.03  -0.01    0.04
#> Fatigue  0.00  0.00   0.03    0.04
# covariate adjustment for an observed context (motif M5)
sim5 <- simulate_ema(N = 40, n_prompts = 30, motifs = "M5", seed = 2)
f5 <- fit_pairs(sim5$data, sim5$vars, covariates = "C")
f5$B_cov
#>                   C
#> NegA     0.60860574
#> PosA     0.12733275
#> Stress   0.02070094
#> Fatigue -0.02540301