A complete pair is two adjacent scheduled prompts of one person (within a
day, when a day column is given) that were both answered. Complete pairs are
the building block of every estimator in the package: fit_pairs,
fit_tilt and fit_ipw regress the states at the
second prompt on the states at the first. This function forms the pairs once
and also collects the prompts whose predecessor was answered, which the
response models of the weighted estimators use.
Arguments
- data
Long data frame with one row per scheduled prompt.
- vars
Names of the state columns.
- id, time
Names of the person and prompt-index columns.
- day
Optional name of a day column; pairs are formed only within a day (the overnight gap is not a lag-1 transition).
- R
Name of the response-indicator column (0/1, no
NA). If the default name is not among the columns ofdata, a prompt counts as answered when allvarsare non-missing; any other name must exist.
Value
A list with y (matrix of states at \(t\)), x
(states at \(t - 1\)), id, time, row_t,
row_tm1 (row indices in data) for the complete pairs;
lag_obs, a list describing every prompt whose predecessor is
answered (its row row_t, its predecessor's row row_tm1, the
predecessor's states x, its own response indicator R,
id and time), which response models on the observed history
use; and vars, n_prompts, response_rate and
Rv (the response indicator in the row order of data).
Details
The time column must be an integer prompt index that
increases by one from one scheduled prompt to the next within a person
(and within a day when day is given); two rows whose indices differ
by more than one are not treated as adjacent. Every scheduled prompt,
answered or not, needs a row.
Examples
sim <- simulate_ema(N = 10, n_prompts = 12, motifs = "M1", seed = 1)
pr <- make_pairs(sim$data, sim$vars)
nrow(pr$y); pr$response_rate
#> [1] 63
#> [1] 0.75
# pairs are formed within days only when a day column is given
simd <- simulate_ema(N = 10, n_prompts = 12, days = 4, seed = 1)
nrow(make_pairs(simd$data, simd$vars)$y)
#> [1] 60
nrow(make_pairs(simd$data, simd$vars, day = "day")$y)
#> [1] 48