Appendix A
R Primer for Longitudinal Analysis
This appendix brings a reader with minimal R to the working baseline the book assumes, and it doubles as the package registry for the whole book. It is not a general R course. Every example uses a longitudinal data shape, drawn from the shipped companion datasets, so that the syntax is learned on the structures the methods chapters actually use: person-period rows, wave-indexed columns, and per-person model fits. A reader fluent in R can skip to the package registry in Section A.8, which is the appendix’s reference core; a reader newer to R should read the sections in order, running each snippet against the companion data as they go.
A.1 Setup and Project Workflow
Install R (version 4.3 or later is assumed throughout) from the Comprehensive R Archive Network, and install RStudio as the integrated development environment; the Quarto publishing system, bundled with recent RStudio, renders the analysis notebooks the book recommends. Work in an RStudio project rather than a bare script folder, because a project fixes the working directory to the project root and lets the here package build paths that survive being moved between machines, which matters for longitudinal work whose pipelines run for hours and are rerun for years. The companion repository has the directory layout of Chapter 36: an immutable data-raw/ of inputs, a built data/ of derived datasets, an R/ of functions, and a figs/ of outputs, with the whole rebuildable from the generation scripts. Download it, open the project file, and restore the package environment before running anything, so that the versions match the ones the book used.
A.2 Objects and Data Essentials
The objects a longitudinal analyst uses most are vectors, factors, data frames, and lists, and each has a longitudinal reading. A vector holds one variable, such as a column of negative-affect scores; a factor is a vector with a fixed set of levels, the natural type for a wave label or an experimental arm, and setting its level order deliberately, rather than accepting the alphabetical default, controls how it enters models and plots. A data frame, or its tidyverse refinement the tibble, holds the person-period table that is the book’s canonical long format, one row per person per occasion, and it is the shape almost every method in the book expects. A list holds heterogeneous objects, and its longitudinal use is to hold one fitted model per person, the structure the idiographic methods of Chapters 24 and 25 build. Reading data uses the readr package for delimited files and the haven package for the SPSS and Stata files that longitudinal datasets often arrive in, both of which return tibbles ready for the verbs of the next section.
A.3 The tidyverse Core Verbs
A dozen verbs do most of the data work in this book, and learning them on longitudinal data is the point of this section. The snippets below run against the shipped affect_ema and school_growth datasets and mirror the idioms of Chapter 5 at a gentler pace. The pipe operator |> passes a result to the next verb, and the grammar reads left to right as a sequence of transformations.
library(dplyr); library(tidyr); library(purrr)
ae <- readRDS("data/affect_ema.rds") # 120 persons x ~84 beeps, long
# select columns, filter rows, create variables
ae |> select(person, day, beep, na, stress) |>
filter(!is.na(na)) |>
mutate(stress_hi = if_else(stress > 3, 1L, 0L))
# group_by + summarise: one row per person (the person-level collapse)
person_means <- ae |>
filter(!is.na(na)) |>
group_by(person) |>
summarise(mean_na = mean(na), sd_na = sd(na), .groups = "drop")
# pivot_wider / pivot_longer: reshape between long and wide
wide <- ae |> filter(person <= 2) |>
select(person, day, beep, na) |>
pivot_wider(names_from = beep, values_from = na, names_prefix = "beep")
long <- wide |> pivot_longer(starts_with("beep"),
names_to = "beep", values_to = "na")
# arrange, across, case_when
ae |> arrange(person, day, beep) |>
mutate(band = case_when(stress > 3 ~ "high",
stress > 2 ~ "mid",
TRUE ~ "low")) |>
group_by(band) |>
summarise(across(c(na, pa), ~ mean(.x, na.rm = TRUE)))
# joins: attach a person-level covariate to the person-period rows
sg <- readRDS("data/school_growth.rds")
roster <- sg |> distinct(child_id, school_id) # one row per child
sg |> left_join(roster, by = "child_id") # attach; anti_join to find orphans
The verbs compose: a within-person centering, the operation the intensive-longitudinal chapters use constantly, is a group_by(person) followed by a mutate that subtracts the person mean, and the recurring hazard, that a person with any missing occasion yields an all-missing group mean, is defeated by na.rm = TRUE in that mean, a point Chapters 31 and 35 return to.
A.4 The Formula Interface and Model Objects
R models are specified by a formula, y ~ x, read as “y is modeled by x,” and the random-effect syntax that the mixed-model chapters use extends it. The dialects differ across engines, and Table A.1 is the correspondence a reader moving between them needs. A fitted model is an object, and the same handful of functions interrogate any of them: summary prints the fit, coef and fixef extract coefficients, ranef extracts random effects, predict produces fitted values, and the broom package’s tidy and augment return the results as tibbles ready for the verbs above.
Table A.1. Random-Effect Formula Syntax Across Engines
| Engine | Random intercept and slope of time within person |
|---|---|
| lme4 | y ~ time + (time | person) |
| nlme | lme(y ~ time, random = ~ time | person) |
| brms | y ~ time + (time | person) (same as lme4; priors added) |
| Mplus | %WITHIN% and %BETWEEN% blocks; s | y ON time; |
Note. The lme4 and brms formulas are identical, which eases the move from frequentist to Bayesian estimation; nlme separates the fixed and random parts; Mplus uses a level-based language. Appendix D gives the full cross-software correspondence for each model family.
A.5 Iteration for Repeated Data
The fit-per-person pattern, central to the idiographic methods of Chapters 24 and 25, splits the data by person and applies a model to each piece, and the purrr package’s map family expresses it cleanly.
library(purrr)
# one lm per person; collect the stress -> na slope for each
slopes <- ae |>
filter(!is.na(na), !is.na(stress)) |>
group_by(person) |>
group_split() |>
map_dbl(~ coef(lm(na ~ stress, data = .x))["stress"])
# slopes is a length-120 vector of person-specific reactivities
The pattern generalizes: map returns a list, map_dbl a numeric vector, and map_dfr a row-bound tibble, so a per-person model’s coefficients, diagnostics, or predictions can be gathered into a tidy frame for the between-person analysis that follows, which is the two-stage structure the book both uses and, in Chapters 16 and 35, cautions against applying without propagating the first-stage uncertainty.
A.6 Graphics Baseline
The book’s figures use ggplot2, whose mental model is a layered grammar: data are mapped to aesthetics, the mapping is realized by geometric layers, and scales, facets, and themes refine the result. A longitudinal spaghetti plot, the workhorse display of Chapter 8, is a few lines.
library(ggplot2); source("R/theme_book_V01.R"); theme_set(theme_book())
ggplot(sg, aes(grade, achievement, group = child_id)) +
geom_line(alpha = 0.1) +
stat_summary(aes(group = 1), fun = mean, geom = "line",
linewidth = 1.2, color = "#B23A48") +
labs(x = "Grade", y = "Achievement")
The theme_book function supplies the house style used throughout, and the palette that matches the book’s figures is defined in the companion; the principle the book presses is to show the data, the thin person lines, and not only the model, the bold mean, which is the visualization obligation of the Chapter 36 reporting checklist.
A.7 Getting Unstuck
Errors are information, and reading them is a skill. Table A.2 decodes the messages a longitudinal analyst meets most. When a problem resists, the reproducible-example discipline, isolating the failure in the smallest self-contained snippet on a shipped dataset, both often reveals the cause and produces the artifact a help forum can act on. The help resources worth knowing are the function documentation reached by ?function, the package vignettes reached by browseVignettes, and the model-specific communities for lme4, lavaan, and the Stan ecosystem.
Table A.2. An Error-Message Decoder
| Message (paraphrased) | Usual cause and fix |
|---|---|
object ’x’ not found | A name is mistyped or the object was never created; check spelling and the pipeline order |
0 (non-NA) cases | Every row is missing on some model variable; a group mean computed without na.rm is the classic culprit |
boundary (singular) fit | A random-effect variance is estimated at zero; simplify the random structure or accept it |
Model failed to converge | Poor scaling or an over-complex model; rescale predictors, try another optimizer, simplify |
contrasts can be applied only to factors with 2 or more levels | A factor predictor has one level after filtering; check the subset |
Note. The second row is the error this book’s own analyses hit most, because the intensive-longitudinal datasets carry missingness in every person, so a within-person centering must use na.rm = TRUE or every group mean becomes missing and no complete cases remain.
A.8 The Master Package Registry
Table A.3 is the appendix’s reference core: every package the book’s analyses use, with the version floor the companion environment pins, the purpose, and the chapters that rely on it. The versions are those under which the book’s results were produced, on R 4.3.3; the companion environment’s lockfile records them exactly, and a reader restoring that environment reproduces the versions rather than merely compatible ones. Table A.4 lists the packages the book discusses but which require a special toolchain, a compiler, an external program, or a sampler, with the setup each needs; these are flagged because their installation is the most common obstacle a reader meets, and their availability should be verified against current documentation, since the sampler and external-program ecosystems move faster than the core packages.
Table A.3. The Master Package Registry (Core, Installable from CRAN)
| Package | Version | Purpose | Chapters |
|---|---|---|---|
| lme4 | 1.1-35 | Mixed models (LMM, GLMM) | 13–16, 23, 34–35 |
| nlme | 3.1-164 | Mixed models; correlation structures | 13–14, 27 |
| lmerTest | 3.1-3 | Denominator df for lme4 | 13–15 |
| mgcv | 1.9-1 | GAMM, TVEM, cyclic smooths | 30, 34–35 |
| lavaan | 0.6-17 | SEM, growth, invariance, RI-CLPM, LCS | 18–21, 34 |
| survival | 3.5-8 | Event-history models | 29 |
| glmnet | 4.1-8 | Regularized regression | 31 |
| rpart | 4.1-23 | Recursive partitioning (trees) | 31 |
| geepack | 1.3-13 | Generalized estimating equations | 12 |
| mice | 3.16-0 | Multiple imputation | 6 |
| MASS | 7.3-60 | mvrnorm; classical models | many |
| Matrix | 1.6-5 | Sparse and dense matrix algebra | 26–27 |
| psych | 2.4-1 | Descriptives, EFA | 3, 18 |
| mvtnorm | 1.2-4 | Multivariate normal density | 16–17 |
| statmod, ordinal, nnet, zoo | — | Gauss–Hermite; ordinal; multinomial; rolling windows | 15–16, 22, 32 |
| ggplot2, dplyr, tidyr, purrr, patchwork | — | Data manipulation and graphics (the tidyverse) | all |
Note. Versions are those under which the book’s results were produced (R 4.3.3); the companion lockfile pins them. A dash groups minor utilities. These packages install from CRAN without a compiler on standard systems.
Table A.4. Packages Requiring a Special Toolchain (Verify at Use)
| Package / tool | Purpose (chapters) | Setup note |
|---|---|---|
| brms, rstan, cmdstanr | Bayesian mixed and dynamic models (17, 25) | Requires a C++ toolchain and the Stan compiler; install CmdStan; long build |
| blavaan | Bayesian SEM (17–18) | Depends on the Stan toolchain above |
| Mplus + MplusAutomation | DSEM, mixtures, invariance (21–22, 25) | Mplus is commercial and external; the R package round-trips input and output |
| lcmm | Latent-class mixed models (22) | Compiles from source; check the current CRAN status |
| dynr, ctsem | Nonlinear and continuous-time dynamics (27, 32) | Compile from source; dynr needs additional system libraries |
| mlVAR, graphicalVAR, qgraph | Network and multilevel VAR (25, 28) | Install as a set; heavy dependency trees |
| renv, targets | Reproducibility stack (36) | Core to the workflow; APIs evolve, so check current docs |
Note. These are the packages whose installation most often blocks a reader. Their availability and interfaces change faster than the core set, so verify each against its current documentation before relying on it; where a book analysis could not run one of these in its environment, the corresponding method was implemented directly and the substitution noted in that chapter’s software box.