Research

Statistical and psychometric methodology for psychological science

The overarching goal of this research program is to strengthen the inferential foundations of psychological science. The work develops, evaluates, and compares statistical models and psychometric methods, asking under what conditions a given method recovers the structures and effects it is meant to recover, and what happens to substantive conclusions when its assumptions fail. Findings are translated into concrete guidance on design, sample size, model selection, and reporting, and into software that makes the methods usable in practice.

Manuscripts currently under review are not listed here. This page describes the directions of the program; see Publications for the published record.

AXIS 1

Nested and dependent data

A long-standing line of work concerns models with discrete latent variables for nested data, especially multilevel latent class and latent transition models. Contributions include sample size recommendations, approaches for incorporating covariates, simultaneous decisions on the number of latent clusters and classes, and analyses of the consequences of ignoring the nesting structure. Related work addresses the specification of random effect structures in linear mixed-effects models, dependency indices such as the intraclass correlation and their robustness under clustering, and information criteria for multilevel mixture models.

Representative work: Park & Yu (2018, Educational and Psychological Measurement); Park & Yu (2018, Structural Equation Modeling); Yu & Park (2014, Multivariate Behavioral Research); Park, Cardwell, & Yu (2020, Methodology); Yu (2013, Behavior Research Methods).

  • Multilevel latent class models
  • Mixed-effects specification
  • ICC and dependency
  • Model selection
  • Sample size planning
Random-coefficient multilevel model as a nested-plate diagram
Random-coefficient model with nested plates for occasions within persons
AXIS 2

Construct representation: factors versus networks

This theme examines the relations between latent variable models and psychometric network models: when the two representations are empirically distinguishable, how partial correlation networks behave when the data are generated by common factors, and how networks can be compared across groups, levels, and populations in a calibrated way. Current work develops unified frameworks for differential network functioning, mixture models that let factor and network structures coexist in a population, and a general account of how the structural form of a latent variable (continuous or categorical, reflective or formative, factor or network) can be chosen on evidence rather than by convention.

Funded by the NSTC project An Alternative Approach of Latent Variable Modeling: Psychometric Network Modeling (2020–2022). Earlier foundations: Anderson & Yu (2007, Psychometrika) on log-multiplicative association models as item response models.

  • Psychometric networks
  • Factor–network equivalence
  • Differential network functioning
  • Latent structure selection
  • IRT and LMA models
Five-node psychometric network with bidirected edges
Gaussian graphical model: edges are the nonzero off-diagonals of the precision matrix
AXIS 3

Within-person dynamics

Work in this area asks how change unfolds inside a person and how research designs constrain what can be learned about it. Topics include the analysis of ecological momentary assessment and other intensive longitudinal data, the characterization of individual change trajectories and their shapes, sampling designs that operate on two time scales, and the comparison of continuous-time and discrete-time modeling frameworks. A parallel strand quantifies and evaluates intervention effects in single-case experimental designs, examining nonoverlap effect size indices and the behavior of inferential methods under serial dependence.

Representative work: Yu (in press, Psychological Methods). Funded by the NSTC projects Exploring Statistical and Methodological Issues in the Analysis of Data from Ecological Momentary Assessment Designs (2025–2027) and Quantifying and Evaluating Intervention Effect in Single-Case Research Design (2022–2023). See also the textbook Analyzing Change.

  • Intensive longitudinal data
  • EMA designs
  • Continuous-time SEM
  • Single-case designs
  • Serial dependence
Random-intercept cross-lagged panel model diagram
Random-intercept cross-lagged panel model separating stable between-person differences from within-person dynamics
AXIS 4

Evidence quality, measurement, and synthesis

This theme addresses the quality of the evidence that psychological research produces and aggregates. In research synthesis, it concerns dependence among effect sizes, selection bias in the published literature, and power planning for moderator inference in sparse meta-regression. In measurement, it spans the psychometric scaling of response anchors in Likert-type scales, scale development and validation, reliability reporting for short forms, applications of aptitude testing to national examinations, and the detection of invalid and insufficient-effort responding in survey data.

Funded by the NSTC projects Investigating Statistical and Methodological Issues in Meta-Analysis: Selection Bias and Dependent Effect Sizes (2023–2024) and Invalid Responses in Survey Data: Impacts and Detection Methods (2024–2025). Representative work: Yu (2022, Survey Research: Method and Application); Yu, Tsai, & Yu (2015, Psychological Testing); Weng, Chen, & Yu (2019, Chinese Journal of Psychology).

  • Meta-analysis
  • Dependent effect sizes
  • Power for moderators
  • Insufficient-effort responding
  • Scale development
Latent class model diagram
Latent class model: a categorical latent variable with mixing proportions as a parameter node
CROSS-CUTTING

Representation and tools

Every axis depends on being able to state a model precisely and communicate it clearly. The umg project develops a unified graphical grammar for statistical models: a small lexicon of nodes (latent and observed, continuous and categorical), edges (stochastic dependence, covariance, deterministic assignment, mixing), and plates (replication and hierarchy) with well-defined semantics, so that a CFA, an IRT model, a multilevel model, a growth mixture model, and a Bayesian hierarchical model can all be drawn in one consistent notation and translated between diagram, model syntax, and code. The grammar is implemented as an R package with TikZ, ggplot2, and DiagrammeR backends. The diagrams on this site are drawn with it.

Software: umg (R; submitted to CRAN); silentema (R; submitted to CRAN); MDLV toolbox (MATLAB; Yu, 2013).

  • Graphical grammar
  • Model visualization
  • R packages
  • Reproducible workflows
The graphical lexicon: node types, edge types, and plates
The lexicon: nodes, edges, and plates

Emerging direction: generative AI and the boundaries of psychological inference

Large language models are entering psychological research as respondents, raters, item writers, and analysts. A recent line of work asks what this does to the inferential conditions of the discipline: when a model-generated response can stand in for a human one, how measurement invariance and validity arguments transfer, and where the boundaries of inference must be redrawn (Yu & Yo, 2026, Chinese Journal of Psychology). Earlier research examined the aggregation of probability judgments from multiple advisors (Budescu & Yu, 2006, 2007).

Funding

Research funding as principal investigator

YearsProjectFunder
2025–2027Exploring Statistical and Methodological Issues in the Analysis of Data from Ecological Momentary Assessment DesignsNSTC 114-2410-H-004-174-MY2
2024–2025Invalid Responses in Survey Data: Impacts and Detection MethodsNSTC 113-2410-H-004-184
2023–2024Investigating Statistical and Methodological Issues in Meta-Analysis: Selection Bias and Dependent Effect SizesNSTC 112-2410-H-004-136
2022–2023Quantifying and Evaluating Intervention Effect in Single-Case Research DesignMOST 111-2410-H-004-168
2020–2022An Alternative Approach of Latent Variable Modeling: Psychometric Network ModelingMOST 109-2410-H-004-075-MY2
2019–2020Methodological Issues in Analyzing Data with Multilevel Structure: Dependency Indices, Sample Sizes, Effect Sizes and PowerMOST 108-2410-H-004-100
2018–2019Covariate Effects in Multilevel Latent Class Models: Comparisons of Different Specifying and Estimating ApproachesMOST 107-2410-H-004-102
2017–2018Multilevel Latent Class Models: Determining the Required Sample Sizes and Developing the Distinguishability and Association MeasuresMOST 106-2410-H-004-064
2010–2015Generalizations and Extensions of Multilevel Latent Transition ModelsNSERC Discovery Grant (Canada)

She has also served as Co-Principal Investigator on funded projects on experience sampling studies of emotion dynamics and depression, the HiTOP model of psychopathology in Taiwan, AI for the public good, and suicide risk modeling for the Taiwan national suicide prevention hotline.