Appendix A

R Primer for Longitudinal Analysis

This appendix brings a reader with minimal R to the working baseline the book assumes, and it doubles as the package registry for the whole book. It is not a general R course. Every example uses a longitudinal data shape, drawn from the shipped companion datasets, so that the syntax is learned on the structures the methods chapters actually use: person-period rows, wave-indexed columns, and per-person model fits. A reader fluent in R can skip to the package registry in Section A.8, which is the appendix’s reference core; a reader newer to R should read the sections in order, running each snippet against the companion data as they go.

A.1 Setup and Project Workflow

Install R (version 4.3 or later is assumed throughout) from the Comprehensive R Archive Network, and install RStudio as the integrated development environment; the Quarto publishing system, bundled with recent RStudio, renders the analysis notebooks the book recommends. Work in an RStudio project rather than a bare script folder, because a project fixes the working directory to the project root and lets the here package build paths that survive being moved between machines, which matters for longitudinal work whose pipelines run for hours and are rerun for years. The companion repository has the directory layout of Chapter 36: an immutable data-raw/ of inputs, a built data/ of derived datasets, an R/ of functions, and a figs/ of outputs, with the whole rebuildable from the generation scripts. Download it, open the project file, and restore the package environment before running anything, so that the versions match the ones the book used.

A.2 Objects and Data Essentials

The objects a longitudinal analyst uses most are vectors, factors, data frames, and lists, and each has a longitudinal reading. A vector holds one variable, such as a column of negative-affect scores; a factor is a vector with a fixed set of levels, the natural type for a wave label or an experimental arm, and setting its level order deliberately, rather than accepting the alphabetical default, controls how it enters models and plots. A data frame, or its tidyverse refinement the tibble, holds the person-period table that is the book’s canonical long format, one row per person per occasion, and it is the shape almost every method in the book expects. A list holds heterogeneous objects, and its longitudinal use is to hold one fitted model per person, the structure the idiographic methods of Chapters 24 and 25 build. Reading data uses the readr package for delimited files and the haven package for the SPSS and Stata files that longitudinal datasets often arrive in, both of which return tibbles ready for the verbs of the next section.

A.3 The tidyverse Core Verbs

A dozen verbs do most of the data work in this book, and learning them on longitudinal data is the point of this section. The snippets below run against the shipped affect_ema and school_growth datasets and mirror the idioms of Chapter 5 at a gentler pace. The pipe operator |> passes a result to the next verb, and the grammar reads left to right as a sequence of transformations.

library(dplyr); library(tidyr); library(purrr)
ae <- readRDS("data/affect_ema.rds")     # 120 persons x ~84 beeps, long

# select columns, filter rows, create variables
ae |> select(person, day, beep, na, stress) |>
      filter(!is.na(na)) |>
      mutate(stress_hi = if_else(stress > 3, 1L, 0L))

# group_by + summarise: one row per person (the person-level collapse)
person_means <- ae |>
  filter(!is.na(na)) |>
  group_by(person) |>
  summarise(mean_na = mean(na), sd_na = sd(na), .groups = "drop")

# pivot_wider / pivot_longer: reshape between long and wide
wide <- ae |> filter(person <= 2) |>
  select(person, day, beep, na) |>
  pivot_wider(names_from = beep, values_from = na, names_prefix = "beep")
long <- wide |> pivot_longer(starts_with("beep"),
                             names_to = "beep", values_to = "na")

# arrange, across, case_when
ae |> arrange(person, day, beep) |>
      mutate(band = case_when(stress > 3 ~ "high",
                              stress > 2 ~ "mid",
                              TRUE       ~ "low")) |>
      group_by(band) |>
      summarise(across(c(na, pa), ~ mean(.x, na.rm = TRUE)))

# joins: attach a person-level covariate to the person-period rows
sg   <- readRDS("data/school_growth.rds")
roster <- sg |> distinct(child_id, school_id)     # one row per child
sg |> left_join(roster, by = "child_id")          # attach; anti_join to find orphans

The verbs compose: a within-person centering, the operation the intensive-longitudinal chapters use constantly, is a group_by(person) followed by a mutate that subtracts the person mean, and the recurring hazard, that a person with any missing occasion yields an all-missing group mean, is defeated by na.rm = TRUE in that mean, a point Chapters 31 and 35 return to.

A.4 The Formula Interface and Model Objects

R models are specified by a formula, y ~ x, read as “y is modeled by x,” and the random-effect syntax that the mixed-model chapters use extends it. The dialects differ across engines, and Table A.1 is the correspondence a reader moving between them needs. A fitted model is an object, and the same handful of functions interrogate any of them: summary prints the fit, coef and fixef extract coefficients, ranef extracts random effects, predict produces fitted values, and the broom package’s tidy and augment return the results as tibbles ready for the verbs above.

Table A.1. Random-Effect Formula Syntax Across Engines

EngineRandom intercept and slope of time within person
lme4y ~ time + (time | person)
nlmelme(y ~ time, random = ~ time | person)
brmsy ~ time + (time | person) (same as lme4; priors added)
Mplus%WITHIN% and %BETWEEN% blocks; s | y ON time;

Note. The lme4 and brms formulas are identical, which eases the move from frequentist to Bayesian estimation; nlme separates the fixed and random parts; Mplus uses a level-based language. Appendix D gives the full cross-software correspondence for each model family.

A.5 Iteration for Repeated Data

The fit-per-person pattern, central to the idiographic methods of Chapters 24 and 25, splits the data by person and applies a model to each piece, and the purrr package’s map family expresses it cleanly.

library(purrr)
# one lm per person; collect the stress -> na slope for each
slopes <- ae |>
  filter(!is.na(na), !is.na(stress)) |>
  group_by(person) |>
  group_split() |>
  map_dbl(~ coef(lm(na ~ stress, data = .x))["stress"])
# slopes is a length-120 vector of person-specific reactivities

The pattern generalizes: map returns a list, map_dbl a numeric vector, and map_dfr a row-bound tibble, so a per-person model’s coefficients, diagnostics, or predictions can be gathered into a tidy frame for the between-person analysis that follows, which is the two-stage structure the book both uses and, in Chapters 16 and 35, cautions against applying without propagating the first-stage uncertainty.

A.6 Graphics Baseline

The book’s figures use ggplot2, whose mental model is a layered grammar: data are mapped to aesthetics, the mapping is realized by geometric layers, and scales, facets, and themes refine the result. A longitudinal spaghetti plot, the workhorse display of Chapter 8, is a few lines.

library(ggplot2); source("R/theme_book_V01.R"); theme_set(theme_book())
ggplot(sg, aes(grade, achievement, group = child_id)) +
  geom_line(alpha = 0.1) +
  stat_summary(aes(group = 1), fun = mean, geom = "line",
               linewidth = 1.2, color = "#B23A48") +
  labs(x = "Grade", y = "Achievement")

The theme_book function supplies the house style used throughout, and the palette that matches the book’s figures is defined in the companion; the principle the book presses is to show the data, the thin person lines, and not only the model, the bold mean, which is the visualization obligation of the Chapter 36 reporting checklist.

A.7 Getting Unstuck

Errors are information, and reading them is a skill. Table A.2 decodes the messages a longitudinal analyst meets most. When a problem resists, the reproducible-example discipline, isolating the failure in the smallest self-contained snippet on a shipped dataset, both often reveals the cause and produces the artifact a help forum can act on. The help resources worth knowing are the function documentation reached by ?function, the package vignettes reached by browseVignettes, and the model-specific communities for lme4, lavaan, and the Stan ecosystem.

Table A.2. An Error-Message Decoder

Message (paraphrased)Usual cause and fix
object ’x’ not foundA name is mistyped or the object was never created; check spelling and the pipeline order
0 (non-NA) casesEvery row is missing on some model variable; a group mean computed without na.rm is the classic culprit
boundary (singular) fitA random-effect variance is estimated at zero; simplify the random structure or accept it
Model failed to convergePoor scaling or an over-complex model; rescale predictors, try another optimizer, simplify
contrasts can be applied only to factors with 2 or more levelsA factor predictor has one level after filtering; check the subset

Note. The second row is the error this book’s own analyses hit most, because the intensive-longitudinal datasets carry missingness in every person, so a within-person centering must use na.rm = TRUE or every group mean becomes missing and no complete cases remain.

A.8 The Master Package Registry

Table A.3 is the appendix’s reference core: every package the book’s analyses use, with the version floor the companion environment pins, the purpose, and the chapters that rely on it. The versions are those under which the book’s results were produced, on R 4.3.3; the companion environment’s lockfile records them exactly, and a reader restoring that environment reproduces the versions rather than merely compatible ones. Table A.4 lists the packages the book discusses but which require a special toolchain, a compiler, an external program, or a sampler, with the setup each needs; these are flagged because their installation is the most common obstacle a reader meets, and their availability should be verified against current documentation, since the sampler and external-program ecosystems move faster than the core packages.

Table A.3. The Master Package Registry (Core, Installable from CRAN)

PackageVersionPurposeChapters
lme41.1-35Mixed models (LMM, GLMM)13–16, 23, 34–35
nlme3.1-164Mixed models; correlation structures13–14, 27
lmerTest3.1-3Denominator df for lme413–15
mgcv1.9-1GAMM, TVEM, cyclic smooths30, 34–35
lavaan0.6-17SEM, growth, invariance, RI-CLPM, LCS18–21, 34
survival3.5-8Event-history models29
glmnet4.1-8Regularized regression31
rpart4.1-23Recursive partitioning (trees)31
geepack1.3-13Generalized estimating equations12
mice3.16-0Multiple imputation6
MASS7.3-60mvrnorm; classical modelsmany
Matrix1.6-5Sparse and dense matrix algebra26–27
psych2.4-1Descriptives, EFA3, 18
mvtnorm1.2-4Multivariate normal density16–17
statmod, ordinal, nnet, zoo—Gauss–Hermite; ordinal; multinomial; rolling windows15–16, 22, 32
ggplot2, dplyr, tidyr, purrr, patchwork—Data manipulation and graphics (the tidyverse)all

Note. Versions are those under which the book’s results were produced (R 4.3.3); the companion lockfile pins them. A dash groups minor utilities. These packages install from CRAN without a compiler on standard systems.

Table A.4. Packages Requiring a Special Toolchain (Verify at Use)

Package / toolPurpose (chapters)Setup note
brms, rstan, cmdstanrBayesian mixed and dynamic models (17, 25)Requires a C++ toolchain and the Stan compiler; install CmdStan; long build
blavaanBayesian SEM (17–18)Depends on the Stan toolchain above
Mplus + MplusAutomationDSEM, mixtures, invariance (21–22, 25)Mplus is commercial and external; the R package round-trips input and output
lcmmLatent-class mixed models (22)Compiles from source; check the current CRAN status
dynr, ctsemNonlinear and continuous-time dynamics (27, 32)Compile from source; dynr needs additional system libraries
mlVAR, graphicalVAR, qgraphNetwork and multilevel VAR (25, 28)Install as a set; heavy dependency trees
renv, targetsReproducibility stack (36)Core to the workflow; APIs evolve, so check current docs

Note. These are the packages whose installation most often blocks a reader. Their availability and interfaces change faster than the core set, so verify each against its current documentation before relying on it; where a book analysis could not run one of these in its environment, the corresponding method was implemented directly and the substitution noted in that chapter’s software box.