Appendix D

Software Crosswalk: R, Mplus, SAS, Stata, and Python

This appendix serves readers embedded in non-R ecosystems and the many longitudinal projects that span several software packages. For each of the book’s core model families it gives the reference model, the R code the book uses, and the corresponding syntax skeletons in Mplus, SAS, Stata, and Python, with the capability caveats that matter. These are orientation skeletons, not tutorials: they show the shape of the command and flag what each package can and cannot do, so that a reader can find the right entry point in a familiar tool. The R code is the code the book runs and has been verified against the shipped data; the non-R skeletons are provided for orientation and should be verified in their native software at use, since syntax and capabilities evolve across versions. The appendix’s highest-value section is the last, the registry of estimation-default differences that silently change results across packages.

D.1 Using This Crosswalk

Each of the following sections names the book’s reference model for a family, gives its R form, and lists the equivalents. The guiding fact is that the same model can carry different default choices in different software, so that fitting “the same model” in two packages can yield different numbers not because either is wrong but because their defaults differ, in the estimator, the degrees of freedom, the standard-error type, or the missing-data handling. Section D.9 catalogues these, and a reader moving a model between packages should consult it before concluding that a discrepancy is an error.

D.2 Linear Mixed Models (Chapters 13–14)

The reference model is a random-intercept-and-slope growth model, \(y_{it} = \beta_0 + \beta_1 t + b_{0i} + b_{1i} t + \varepsilon_{it}\), with the random effects bivariate normal.

# R (lme4)      : lmer(y ~ time + (time | person), data = d)
# R (nlme)      : lme(y ~ time, random = ~ time | person, data = d)
# Mplus         : %WITHIN% s | y ON time;  %BETWEEN% y s;  (TYPE = TWOLEVEL RANDOM)
# SAS           : PROC MIXED; CLASS person; MODEL y = time / SOLUTION DDFM=KR;
#                 RANDOM intercept time / SUBJECT=person TYPE=UN;
# Stata         : mixed y time || person: time, cov(unstructured) reml
# Python        : statsmodels MixedLM.from_formula("y ~ time",
#                 groups="person", re_formula="~time")

Capability caveats: lme4 does not report denominator degrees of freedom by default (the lmerTest package adds Satterthwaite or Kenward-Roger), whereas SAS and Stata offer Kenward-Roger directly; statsmodels fits the model but its random-effects covariance options and small-sample inference are more limited than the others.

D.3 Generalized Linear Mixed Models (Chapter 15)

# R           : glmer(y ~ time + (time|person), family=poisson, data=d)
# R (alt)     : glmmTMB(y ~ time + (time|person), family=nbinom2, data=d)
# Mplus       : as TWOLEVEL RANDOM with CATEGORICAL or COUNT ARE y;
# SAS         : PROC GLIMMIX; MODEL y = time / DIST=POISSON SOLUTION;
#               RANDOM intercept time / SUBJECT=person;
# Stata       : mepoisson y time || person: time   (or melogit for binary)
# Python      : statsmodels / PyMC (Bayesian) for non-Gaussian mixed models

Capability caveats: the crucial difference is the approximation to the integral over the random effects. lme4 and SAS GLIMMIX default to Laplace or adaptive Gauss-Hermite quadrature; a package that uses penalized quasi-likelihood instead can bias variance components for binary outcomes, a difference that changes results, not only speed.

D.4 SEM, Growth Curves, and Invariance (Chapters 18–20)

# R (lavaan) latent growth:
#   model <- 'i =~ 1*y1 + 1*y2 + 1*y3 + 1*y4
#             s =~ 0*y1 + 1*y2 + 2*y3 + 3*y4'
#   growth(model, data = d)
# Mplus       : i s | y1@0 y2@1 y3@2 y4@3;   (MODEL command)
# SAS         : PROC CALIS; (LINEQS / path syntax)
# Stata       : sem (i -> y1@1 ...) (s -> ...) , means(i s)
# Python      : semopy (model syntax similar to lavaan; verify status)

Capability caveats: measurement-invariance testing across waves or groups is well supported in lavaan (with the semTools helpers), Mplus, and Stata; the categorical-indicator estimators (weighted least squares) and their fit statistics differ in default across packages, so a fit index compared across engines must use the same estimator.

D.5 The Random-Intercept Cross-Lagged Panel Model (Chapter 21)

# R (lavaan): random intercepts RIx =~ 1*x1 + 1*x2 + ...; RIy =~ 1*y1 + ...
#   within components wx1..wy4; regress wx2 ~ wx1 + wy1; wy2 ~ wy1 + wx1; etc.
#   fix observed residual variances to 0 for single indicators.
# Mplus     : the widely circulated RI-CLPM template (Hamaker) translates directly;
#             random intercepts as BETWEEN-level factors with unit loadings.

Capability caveats: the model is most commonly fit in lavaan and Mplus, whose templates are near-identical; the single-indicator decomposition (fixing measurement error to zero) is the same device in both, and the multi-group extension for moderation follows the general SEM group syntax of Section D.4.

D.6 Growth Mixture and Latent-Class Models (Chapter 22)

# R (lcmm)  : hlme(y ~ time, random=~time, subject="person", ng=3, mixture=~time)
# Mplus     : TYPE = MIXTURE; CLASSES = c(3); the field's reference implementation
# Stata     : gsem with latent class (limited); or the traj-style approach
# SAS       : PROC TRAJ (group-based trajectory modeling lineage)

Capability caveats: this is the one family where the honest statement is that the software is Mplus-centric. Mplus’s mixture facilities, including the enumeration statistics and the auxiliary-variable machinery, are more complete and more widely used than the alternatives; lcmm is the strongest R option, and the SAS trajectory procedure encodes a specific (group-based) modeling tradition rather than the general growth mixture model.

D.7 Survival and Event History (Chapter 29)

# R (survival): coxph(Surv(time, event) ~ x, data = d)
#               survfit(Surv(time,event) ~ 1) for Kaplan-Meier
# SAS         : PROC PHREG; MODEL time*event(0) = x;  (PROC LIFETEST for KM)
# Stata       : stset time, failure(event); stcox x
# Python      : lifelines CoxPHFitter().fit(d, "time", "event")

Capability caveats: the discrete-time survival approach the book emphasizes, fitting a complementary-log-log or logistic model to a person-period expansion, is available in every package as an ordinary generalized linear model on the expanded data, which is often the most portable route across ecosystems.

D.8 Dynamic, Continuous-Time, and Other Families (Chapters 12, 25–27, 30)

The dynamic-structural-equation and continuous-time families are the capability frontier, and here the packages diverge most. Multilevel vector-autoregression is fit in R with mlVAR and Bayesian dynamic structural equation modeling in Mplus with its DSEM facility, which is the reference implementation; continuous-time models are fit in R with ctsem or dynr, and bespoke dynamics in Stan. Generalized estimating equations (Chapter 12) are fit with R’s geepack, SAS PROC GENMOD with a REPEATED statement, and Stata’s xtgee. Generalized additive mixed models (Chapter 30) are an R strength through mgcv, with partial equivalents elsewhere. For these frontier families the capability differences are large enough that the choice of software is often dictated by the method rather than by preference, and Section D.9’s caution about differing defaults applies with special force.

D.9 The Estimation-Default Gotcha Registry

Table D.1 is the appendix’s most valuable content: the estimation defaults that differ across packages and thereby change results when the “same” model is fit in two of them. A discrepancy between packages should be checked against this table before it is treated as an error, because in most cases it is a difference of default, not of correctness.

Table D.1. Estimation-Default Differences That Change Results

IssueThe differenceConsequence
REML vs. MLlme4/SAS/Stata default to REML for linear mixed modelsVariance estimates differ; likelihood-ratio tests of fixed effects require ML
Denominator dfKenward-Roger in SAS/Stata; absent by default in lme4Fixed-effect \(p\)-values differ, especially in small samples
GLMM integrationLaplace / adaptive quadrature vs. penalized quasi-likelihoodVariance components for binary outcomes can be biased under PQL
SEM estimatorMaximum likelihood vs. robust ML vs. weighted least squares for categorical indicatorsFit indices and standard errors are not comparable across estimators
Missing-data defaultFull-information ML in some SEM software; listwise deletion in some proceduresDifferent samples analyzed; biased estimates under listwise deletion
Standard-error typeModel-based vs. robust (sandwich) defaultsInference differs under misspecification or clustering

Note. Each row is a default, not a mistake, and each can make two packages disagree on the same model. The remedy is to set the estimator, the degrees-of-freedom method, the integration, and the missing-data handling explicitly and identically before comparing, rather than to trust that “the same model” means the same computation.

D.10 File Exchange and When to Switch

Moving data between ecosystems uses R’s haven package, which reads and writes the SPSS, Stata, and SAS transport formats, and the MplusAutomation package, which writes Mplus input, runs the model, and reads the output back into R, enabling a hybrid workflow in which R prepares the data and post-processes the results while Mplus fits a model only it can fit. The decision to switch software is, per family, a decision the capability caveats above make for you: stay in R for mixed models, flexible models, and survival, where it is as capable as any package and better integrated with the reproducibility stack of Chapter 36; reach for Mplus for growth mixtures and Bayesian dynamic structural equation modeling, where it is the reference implementation; and use the specialized R packages for continuous-time and network dynamics. The portable default, when a model can be expressed as a generalized linear model on an expanded dataset, as discrete-time survival and some dynamic models can, is to use that expression, because it runs in every package and travels without translation.