Automatic item generation produces items faster than pretesting can calibrate them. coldstart predicts each new item’s difficulty from its features, uses the prediction as a robust prior, and finds template families where the prediction cannot be trusted.
Legacy items have calibrated difficulties; new generated items have only features (any numeric representation works, such as embeddings from any text model or cognitive-attribute codes) and a template family. In the simulation, one family’s template drifted, making its new items harder than its history suggests.
library(coldstart)
sim <- cs_simulate(n_train = 400, n_new = 150, seed = 3)
it <- sim$items
tr <- it$set == "train"
table(it$set)
#>
#> new train
#> 150 400
pr <- cs_predictor(it$b_legacy[tr], sim$features[tr, ], it$family[tr], seed = 1)
pr
#> <cs_predictor> ridge, lambda = 1 | 400 legacy items | 10 families
#> predictive SD: 0.548 (seen family), 0.788 (unseen family)
pred <- predict(pr, sim$features[!tr, ], it$family[!tr])
head(pred)
#> item family mean sd family_seen
#> G0401 G0401 F07 1.3051977 0.5480338 TRUE
#> G0402 G0402 F08 -1.8811880 0.5480338 TRUE
#> G0403 G0403 F11 -0.4888554 0.7876555 FALSE
#> G0404 G0404 F09 -0.5278875 0.5480338 TRUE
#> G0405 G0405 F06 -0.1683571 0.5480338 TRUE
#> G0406 G0406 F02 -0.4489074 0.5480338 TRUE
The predictive SD is estimated out of sample. It is larger for a family never seen in training, because the family’s own effect is then unknown.
plan <- cs_plan(pred, target_sd = 0.3)
summary(plan[c("n_with_prior", "n_without_prior")])
#> n_with_prior n_without_prior
#> Min. :38.00 Min. : 54.00
#> 1st Qu.:39.00 1st Qu.: 55.00
#> Median :41.00 Median : 58.00
#> Mean :46.24 Mean : 64.15
#> 3rd Qu.:48.75 3rd Qu.: 64.75
#> Max. :94.00 Max. :134.00
resp <- cs_responses(sim, n_per_item = 25, seed = 4)
cal <- cs_calibrate(resp, pred)
truth <- it$b_true[match(cal$item, it$item)]
c(baseline = sqrt(mean((cal$base_mean - truth)^2)),
with_prior = sqrt(mean((cal$post_mean - truth)^2)))
#> baseline with_prior
#> 0.5752105 0.3839569
fam <- setNames(it$family[!tr], it$item[!tr])
chk <- cs_check(cs_calibrate(cs_responses(sim, 100, seed = 5), pred), fam)
chk
#> family n_items mean_z rms_z coverage p_value trustworthy
#> 1 F01 13 1.91170195 2.2174671 0.3076923 1.033704e-08 FALSE
#> 9 F09 18 0.26358231 1.1677767 0.7777778 1.379219e-01 TRUE
#> 7 F07 10 -0.08620829 1.2031628 0.9000000 1.523653e-01 TRUE
#> 6 F06 19 -0.03607136 1.0696458 0.8947368 2.974532e-01 TRUE
#> 4 F04 7 0.23966538 1.0288285 0.8571429 3.875306e-01 TRUE
#> 3 F03 12 0.19518559 1.0150236 0.8333333 4.169631e-01 TRUE
#> 8 F08 13 0.23874499 1.0076796 0.8461538 4.324513e-01 TRUE
#> 2 F02 9 -0.10144038 0.8624153 0.8888889 6.689603e-01 TRUE
#> 5 F05 15 0.32898050 0.8516693 0.9333333 7.610460e-01 TRUE
#> 11 F11 19 -0.31884214 0.7773376 1.0000000 9.066038e-01 TRUE
#> 10 F10 15 -0.08314448 0.7429080 1.0000000 9.121243e-01 TRUE
unique(it$family[it$rogue])
#> [1] "F01"
Priors for families that fail the check are withdrawn before the final calibration:
resp100 <- cs_responses(sim, 100, seed = 5)
final <- cs_calibrate(resp100, cs_distrust(pred, chk))
head(final[c("item", "n", "post_mean", "post_sd")])
#> item n post_mean post_sd
#> 1 G0401 100 0.7082320 0.2175301
#> 2 G0402 100 -1.3150924 0.2362350
#> 3 G0403 100 -1.1280780 0.2314078
#> 4 G0404 100 0.6417285 0.2345004
#> 5 G0405 100 -0.7933419 0.2231148
#> 6 G0406 100 -0.6901677 0.2011189