Appendix C

Book Datasets and Codebooks

This appendix documents the datasets the book uses, and it is the reproducibility backbone of the whole project. Nearly every worked example draws on a small set of shipped, simulated datasets whose true data-generating parameters are known and published here, because the book’s pedagogy rests on the ability to check an estimate against the truth that produced it. This appendix gives the master catalogue, the known-truth parameter tables for the flagship datasets, the pointers to the external datasets the book references, and the instructions for regenerating every simulated dataset from its script. The full variable-level codebooks ship in the companion repository under Examples/codebooks/; this appendix is the reference summary and the truth registry.

C.1 The Known-Truth Design Philosophy

Almost all of the book’s datasets are simulated, and this is a pedagogical choice, not a convenience. When the data-generating process is known, an estimate can be compared to the quantity it targets, so that a method is demonstrated to recover a truth rather than merely to produce a plausible number, and a reader can see not only how to run an analysis but whether the analysis works. Every worked example in the methods chapters is an audit in this sense: the growth model recovers the planted growth, the cross-lagged model recovers the planted cross-lag, the survival model recovers the planted hazard, and where the recovery is imperfect, the imperfection is itself instructive, as when a short series biases an autoregressive estimate downward or a plug-in variability index attenuates a second-stage effect. The simulated datasets are made realistic by sourcing their parameter values from the ranges the empirical literature reports, so that a diary study’s autocorrelation, a trial’s effect size, and an aging cohort’s decline rate are of plausible magnitude; each such sourcing claim is one a careful reader should verify against the cited literature, and the book flags them as design choices rather than empirical findings. The known-truth parameters are stored with each dataset as an attribute, retrievable in R by attr(dataset, "truth"), so that the truth travels with the data and no analysis need rely on a separately maintained record that could drift from the object.

C.2 The Master Dataset Catalogue

Table C.1 lists the shipped datasets with their design, the chapters that use them, and the generation script and seed that produce them. Every dataset is rebuildable exactly from its script and seed, and the version suffix on each script follows the lab’s file-versioning standard.

Table C.1. The Shipped Dataset Catalogue

DatasetDesignChaptersScript (seed)
affect_ema120 persons \(\times\) 84 beeps EMA; VAR(1)1, 3, 25, 31, 33, 35, 36gen_affect_ema (20260722)
school_growth1200 children, 3 cohorts, grades 3–81, 2, 14, 18–21, 34gen_school_growth (20260722)
therapy_rct240 patients, 2 arms, 12 weeks2, 33, 36gen_therapy_rct (20260722)
daily_couples100 couples \(\times\) 21 days, 2 partners9, 23, 25gen_daily_couples (20260722)
n1_mood1 subject \(\times\) many days24gen_n1_mood (20260722)
panel_simmulti-wave panel21gen_panel_sim (20260723)
classes_simmixture of trajectory classes22gen_classes_sim (20260725)
dfm_simdynamic-factor series26gen_dfm_sim (20260732)
ct_simbivariate continuous-time (OU)27gen_ct_sim (20260801)
symptom_netmultivariate symptom network28gen_symptom_net (20260803)
relapsediscrete-time survival, competing risks29gen_relapse (20260729)
growth_gamnonlinear growth spurt30gen_flexcurves (20260810)
growth_subgroupsgrowth by a covariate tree31gen_growth_subgroups (20260814)
bistable_simfold-bifurcation SDE32gen_bistable (20260816)
mrt_emamicro-randomized trial33gen_mrt_ema (20260818)
achieve_accelaccelerated cohort, 3 scalings34gen_achieve_accel (20260820)
wiscbivariate cascade, RI-CLPM truth34gen_wisc (20260822)
burst_agingmeasurement-burst aging34gen_burst_aging (20260824)
couples_diarydyadic APIM diary35gen_couples_diary (20260826)
work_weekworker diary, weekly cycle35gen_work_week (20260828)

Note. Each row is rebuildable from its script and seed. Chapter-specific datasets (the lower rows) were built where a planted structure the shipped core lacked was needed to audit a method; the reused core datasets (upper rows) carry several chapters each. Full variable-level codebooks are in Examples/codebooks/.

C.3 Known-Truth Parameter Tables

The flagship datasets’ true parameters are given here so that a reader auditing an analysis has the target values to hand. Table C.2 gives the core recurring datasets; Table C.3 gives representative chapter-specific ones. These are the values stored in each dataset’s truth attribute.

Table C.2. Known-Truth Parameters: Core Datasets

DatasetTrue parameters
affect_emaWithin-person VAR(1): autoregression of negative affect \(\approx 0.35\); contemporaneous stress-to-negative-affect reactivity \(\approx 0.35\) (person-varying, SD \(\approx 0.08\)); person means SD \(\approx 0.44\); missing-at-random beeps
school_growthGrade-centered growth \(185 + 12(\text{grade}-3) - 0.7(\text{grade}-3)^2\); between-child intercept SD 20, slope SD 3.0, correlation \(-0.20\); between-school SD 8; residual SD 9
therapy_rctWeekly depression decline steeper in the treatment arm; arm-by-week effect \(\approx -0.5\) per week; symptom-dependent (informative) dropout

Note. These are the datasets reused across the most chapters, so their truth is quoted most often. The affect_ema reactivity of about 0.35 is the exemplar recovered in Chapters 35 and 36; the school_growth decelerating curve anchors Chapters 1, 14, and 34.

Table C.3. Known-Truth Parameters: Selected Chapter-Specific Datasets

DatasetTrue parameters
bistable_simDouble-well SDE \(dx = (-x^3 + x + c)\,dt + 0.18\,dW\); control \(c\) ramps \(-0.60 \to 0.95\); tipping near step 497; early-warning signals (rising autocorrelation and SD) lead the tip
achieve_accelLatent decelerating growth; SES on level \(0.50\) and on gain \(0.15\); three monotone vertical scalings of one latent (linear, convex, concave)
wiscRI-CLPM: correlated stable traits (\(r = 0.50\)); within-person cross-lags verbal-to-performance \(0.25\), performance-to-verbal \(0.05\); SES moderation \(0.12\)
burst_agingBurst-level mean reaction time \(650/658/705\); within-person SD \(55/82/108\) (leads the mean); autocorrelation \(0.25/0.36/0.47\); informative attrition before burst 3
couples_diaryAPIM within-person actor \(0.40\), partner \(0.20\); between-person support effect \(0.60\); asymmetric cross-partner lag \(0.25\) vs. \(0.10\); satisfaction moderation \(0.15\)
work_weekWeekly vigor cycle (mid-week trough, weekend peak); weekend effect \(0.50\); evening-detachment-to-next-morning-vigor recovery \(0.35\)

Note. The full parameter set for each dataset, including variance components and seeds, is in the dataset’s truth attribute and its codebook. These are the targets the corresponding chapters’ analyses recover, and quoting them here lets a reader reproduce the audit independently.

C.4 External and Package Datasets

The book references a small number of datasets it does not ship, and Table C.4 gives their sources and access notes. These are classic teaching datasets whose availability and licensing should be verified at use, since package contents and repository access change over time. Where the book demonstrates a method on one of these, it also provides a shipped simulated analogue, so that a reader without access to the external data can still run the example.

Table C.4. External Datasets Referenced by the Book

DatasetSourceAccess note
sleepstudylme4 packageShips with the package; the canonical random-slope example
riesbyLongitudinal-methods teaching corpusWidely circulated depression-trial data; verify the current source and terms
alcohol_useSinger & Willett teaching dataFrom the applied-longitudinal-analysis materials; verify access

Note. These external datasets are referenced for continuity with the longitudinal-methods literature. Their access and licensing are not guaranteed stable, so each should be verified at use; the book’s shipped simulated datasets are the primary, self-contained teaching material.

C.5 Regeneration and Versioning

Every simulated dataset is rebuildable from its generation script, which carries the lab-standard header, sets its seed explicitly, and writes both an .rds and a .csv with an accompanying codebook. To rebuild the full set, source the generation scripts in the companion’s Examples/R/ directory; each is self-contained and depends only on base R and a small set of core packages. Datasets are versioned like scripts, with a version suffix that increments on any change to the generating process, so that a result can always be traced to the exact dataset version that produced it, and a change to a data-generating process never silently alters a downstream result. This single-source-of-truth discipline, the codebook and the truth attribute both generated from the same script that makes the data, is what keeps the documentation from drifting away from the object it documents, and it is the dataset-level instance of the reproducibility doctrine Chapter 36 sets out for the project as a whole.