Package {driftwatch}


Title: Sequential Monitoring of Item Parameter Drift
Version: 0.1.0
Description: Ongoing surveillance of item parameter drift for continuous testing programs and pre-equated item banks. Estimates item difficulty in rolling calibration windows, runs sequential cumulative sum (CUSUM) detection (Page, 1954, <doi:10.1093/biomet/41.1-2.100>; applied to testing by Veerkamp and Glas, 2000, <doi:10.3102/10769986025004373>) with change-point estimation, separates gradual drift from abrupt jumps, tunes alarm thresholds for a bank-wide false-alarm target by simulation on the program's own design, quantifies score and pass-rate impact, and records recommended actions in an audit log.
License: MIT + file LICENSE
URL: https://github.com/edidatasolutions/driftwatch, https://edidatasolutions.github.io/driftwatch/
BugReports: https://github.com/edidatasolutions/driftwatch/issues
Encoding: UTF-8
Depends: R (≥ 4.1)
Imports: stats, utils
Suggests: knitr, markdown
VignetteBuilder: knitr
Config/roxygen2/version: 8.1.0
NeedsCompilation: no
Packaged: 2026-09-28 00:37:23 UTC; User
Author: Daniel Edi ORCID iD [aut, cre, cph]
Maintainer: Daniel Edi <danieledi2026@gmail.com>
Repository: CRAN
Date/Publication: 2026-10-07 09:50:16 UTC

Recommended actions and audit log

Description

Turns monitoring alarms into recommended actions with a written rationale:

Usage

dw_actions(
  monitor,
  anchors = character(0),
  retire_at = 0.5,
  analyst = NA_character_
)

Arguments

monitor

A 'dw_monitor'.

anchors

Item ids used as equating anchors.

retire_at

Abrupt magnitude (logits) that triggers retirement.

analyst

Name recorded in the log.

Value

Audit-log data frame, one row per alarmed item.

Examples

sim <- dw_simulate(n_items = 60, n_windows = 20, mean_n = 60,
                   onset_range = c(5, 12), seed = 1)
mon <- dw_monitor(dw_estimate(sim$responses, sim$bank), h = 8)
log <- dw_actions(mon, anchors = sim$bank$item[1:10], analyst = "DE")
if (nrow(log)) log[, c("item", "type", "magnitude", "action")]

Window-level item difficulty estimates

Description

Rasch difficulty for every item x window, with examinee ability treated as known (from operational scoring on the rest of the form). Estimation is penalized maximum likelihood with a weak N(b_bank, 'prior_sd'^2) penalty that only matters for all-correct or all-incorrect windows. Newton steps run for all item-windows at once.

Usage

dw_estimate(responses, bank, prior_sd = 3)

Arguments

responses

Long data frame: 'window' (integer), 'item', 'theta', 'x'.

bank

Reference parameters: 'item', 'b', and optionally 'se'.

prior_sd

SD of the weak penalty.

Value

A 'dw_estimates' object: matrices 'b_hat', 'se', 'n' and 'z' (items x windows; 'z' is the standardized deviation from the bank, NA where an item was not administered), plus 'bank' and 'responses'.

Examples

sim <- dw_simulate(n_items = 60, n_windows = 20, mean_n = 60,
                   onset_range = c(5, 12), seed = 1)
est <- dw_estimate(sim$responses, sim$bank)
round(est$z[1:5, 1:6], 2)

Score and pass-rate impact of drifted items

Description

For a pre-equated form scored by true-score conversion (the raw score is mapped to theta through the test characteristic curve of the scoring parameters), an examinee of ability theta is reported at 'TCC_scoring^-1(TCC_current(theta))'. The function compares scoring scenarios against the current (true or best-estimate) difficulties:

keep

Score with banked parameters for every item.

remove

Drop flagged items from the form; score the rest with banked parameters.

recalibrate

Score flagged items with their current estimates.

and reports reported-score bias at the cut and the pass rate in the population against the correct pass rate.

Usage

dw_impact(
  form,
  bank,
  current,
  flagged,
  recalibrated = NULL,
  cut = 0,
  theta_mean = 0,
  theta_sd = 1
)

Arguments

form

Item ids on the form.

bank

Banked parameters ('item', 'b').

current

Named vector of current difficulties: the truth in a simulation, or the latest window estimates in practice.

flagged

Item ids flagged by monitoring.

recalibrated

Named vector of re-estimated difficulties for flagged items (default: 'current[flagged]').

cut

Passing standard on the theta scale.

theta_mean, theta_sd

Examinee population.

Value

Data frame: 'scenario', 'n_items', 'bias_at_cut', 'mean_abs_bias', 'pass_rate', 'pass_rate_error' (vs the correct rate).

Examples

# A 20-item form where one item became 0.8 logits harder
current <- setNames(c(0.8, rep(0, 19)), paste0("q", 1:20))
bank <- data.frame(item = names(current), b = 0)
dw_impact(names(current), bank, current, flagged = "q1", cut = 0)

Sequential drift monitoring

Description

Runs a two-sided CUSUM on each item's standardized deviations 'z_t = (b_t - b_bank) / SE': 'S+[t] = max(0, S+[t-1] + z[t] - k)' and likewise downward. An alarm is raised the first window either statistic exceeds 'h'. Drift type is then classified by comparing step and ramp fits to the difficulty series (profiling the change point), using data up to 'followup' windows after the alarm. 'change_point' is the onset under the better-fitting model; 'cusum_change_point' is the standard CUSUM estimator (the window after the statistic last left zero), which locates abrupt changes well but lags gradual onsets.

Usage

dw_monitor(estimates, h, k = 0.5, followup = 3, min_log_lr = 1)

Arguments

estimates

A 'dw_estimates' object.

h

Alarm threshold, ideally from [dw_tune()].

k

Reference value (half the shift, in SE units, the chart is tuned to detect).

followup

Windows after the alarm used to classify the change ('Inf' = all data to date).

min_log_lr

Minimum log likelihood ratio between the step and ramp models for a type to be assigned; weaker evidence gives '"undetermined"'. Soon after an alarm a ramp and a step often cannot be told apart; rerun with more windows to resolve them.

Value

A 'dw_monitor' object; '$items' has one row per item: 'alarm' (logical), 'alarm_window', 'direction' ('harder' / 'easier'), 'change_point', 'type' ('abrupt' / 'gradual' / 'undetermined'), 'type_log_lr' (evidence for the better model), 'magnitude' (current displacement under the better model, logits), 'rate' (logits per window, gradual only).

Examples

sim <- dw_simulate(n_items = 60, n_windows = 20, mean_n = 60,
                   onset_range = c(5, 12), seed = 1)
est <- dw_estimate(sim$responses, sim$bank)
mon <- dw_monitor(est, h = 8)
mon
table(alarm = mon$items$alarm, truth = sim$truth$type)

Simulate a continuously administered item bank with known drift

Description

Each window, every item is answered by a Poisson number of examinees whose abilities are known from operational scoring (the population mean may trend over time; this is not drift). Items are stable, drift gradually (linear from an onset window) or jump abruptly at an onset window.

Usage

dw_simulate(
  n_items = 300,
  n_windows = 40,
  mean_n = 80,
  p_gradual = 0.1,
  p_abrupt = 0.05,
  slope_range = c(0.02, 0.06),
  jump_range = c(0.4, 1),
  onset_range = c(5, 30),
  theta_trend = 0.01,
  ref_se = 0.05,
  seed = NULL
)

Arguments

n_items, n_windows

Bank size and number of windows.

mean_n

Mean responses per item per window.

p_gradual, p_abrupt

Share of items with each drift type.

slope_range

Absolute gradual slope per window (logits).

jump_range

Absolute abrupt jump (logits).

onset_range

Windows in which drift can begin.

theta_trend

Change in examinee mean ability per window.

ref_se

Standard error of the banked (reference) difficulties.

seed

Optional seed.

Value

A 'dw_sim': '$responses' ('window', 'item', 'theta', 'x'), '$bank' ('item', 'b', 'se'), '$truth' ('item', 'type', 'onset', 'size', 'b_true_final') and '$b_path' (items x windows matrix of true difficulty).

Examples

sim <- dw_simulate(n_items = 60, n_windows = 20, mean_n = 60,
                   onset_range = c(5, 12), seed = 1)
table(sim$truth$type)

Tune the alarm threshold for a bank-wide false-alarm target

Description

An item raises a false alarm over the monitoring horizon exactly when the maximum of its CUSUM statistics exceeds 'h'. So 'h' is the '1 - target' quantile of that maximum under no drift. With ‘method = "design"', the null is simulated on the program’s own design: the same items, windows, examinee abilities and sample sizes, with responses regenerated from the banked difficulties, then re-estimated and re-standardized exactly as in monitoring. This captures small-sample non-normality of 'z' and sparse windows. 'method = "normal"' treats 'z' as iid N(0, 1), which is fast and useful for planning a bank that does not exist yet.

Usage

dw_tune(
  estimates = NULL,
  target = 0.01,
  k = 0.5,
  method = c("design", "normal"),
  n_rep = 20,
  n_windows = NULL,
  seed = NULL
)

Arguments

estimates

A 'dw_estimates' object (required for '"design"').

target

Probability that a non-drifting item alarms at least once over the horizon. Expected false alarms for the bank = 'target * n_items'.

k

Reference value (as in [dw_monitor()]).

method

'"design"' or '"normal"'.

n_rep

Null replicates of the whole bank ('"design"') or simulated item series ('"normal"').

n_windows

Horizon for ‘"normal"' (default: the estimates’ windows).

seed

Optional seed.

Value

A list: 'h', 'target', 'k', 'method', 'expected_false_alarms' (per bank, when estimates are given), and 'null_max' (the simulated maxima).

Examples

sim <- dw_simulate(n_items = 60, n_windows = 20, mean_n = 60,
                   onset_range = c(5, 12), seed = 1)
est <- dw_estimate(sim$responses, sim$bank)
dw_tune(est, target = 0.02, method = "normal", seed = 1)$h

# Design-based tuning (recommended) simulates the whole bank n_rep times.
dw_tune(est, target = 0.02, n_rep = 5, seed = 1)$h


Two-point drift check (the conventional baseline)

Description

Pools the first and last blocks of windows, estimates each item's difficulty in both, and flags items whose robust z of the difference (median/MAD standardized) exceeds 'crit': the usual displacement check at equating time.

Usage

dw_twopoint(estimates, early = 1:5, late = NULL, crit = 2.7)

Arguments

estimates

A 'dw_estimates' object.

early, late

Window indices forming the two calibrations.

crit

Robust-z criterion.

Value

Data frame: 'item', 'd', 'robust_z', 'flag'.

Examples

sim <- dw_simulate(n_items = 60, n_windows = 20, mean_n = 60,
                   onset_range = c(5, 12), seed = 1)
tp <- dw_twopoint(dw_estimate(sim$responses, sim$bank))
table(flag = tp$flag, truth = sim$truth$type)