Skip to contents

A complete pair is two adjacent scheduled prompts of one person (within a day, when a day column is given) that were both answered. Complete pairs are the building block of every estimator in the package: fit_pairs, fit_tilt and fit_ipw regress the states at the second prompt on the states at the first. This function forms the pairs once and also collects the prompts whose predecessor was answered, which the response models of the weighted estimators use.

Usage

make_pairs(data, vars, id = "id", time = "time", day = NULL, R = "R")

Arguments

data

Long data frame with one row per scheduled prompt.

vars

Names of the state columns.

id, time

Names of the person and prompt-index columns.

day

Optional name of a day column; pairs are formed only within a day (the overnight gap is not a lag-1 transition).

R

Name of the response-indicator column (0/1, no NA). If the default name is not among the columns of data, a prompt counts as answered when all vars are non-missing; any other name must exist.

Value

A list with y (matrix of states at \(t\)), x

(states at \(t - 1\)), id, time, row_t,

row_tm1 (row indices in data) for the complete pairs;

lag_obs, a list describing every prompt whose predecessor is answered (its row row_t, its predecessor's row row_tm1, the predecessor's states x, its own response indicator R,

id and time), which response models on the observed history use; and vars, n_prompts, response_rate and

Rv (the response indicator in the row order of data).

Details

The time column must be an integer prompt index that increases by one from one scheduled prompt to the next within a person (and within a day when day is given); two rows whose indices differ by more than one are not treated as adjacent. Every scheduled prompt, answered or not, needs a row.

Examples

sim <- simulate_ema(N = 10, n_prompts = 12, motifs = "M1", seed = 1)
pr <- make_pairs(sim$data, sim$vars)
nrow(pr$y); pr$response_rate
#> [1] 63
#> [1] 0.75
# pairs are formed within days only when a day column is given
simd <- simulate_ema(N = 10, n_prompts = 12, days = 4, seed = 1)
nrow(make_pairs(simd$data, simd$vars)$y)
#> [1] 60
nrow(make_pairs(simd$data, simd$vars, day = "day")$y)
#> [1] 48