AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Evaluating Model Retraining under Drift: Paired Comparisons of Cumulative Subgroup Disparity

arXiv · AI, language, vision and robotics · article · Sep 9, 2026 · UTC

Choosing when to retrain a deployed classifier requires assessing subgroup error rates across the sequence of models used, including periods between updates. We compare complete scheduled, loss-triggered, and subgroup-gap-triggered policies with retaining the initial model on the same observations and delayed labels. For true-positive and false-positive rates separately, the outcome is the paired difference in absolute subgroup gaps summed over deployment windows. Population evaluation in simulation, action records, and alternative schedules assess how measurement and retraining behaviour affe

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T19:52:05.078Z. This is not the publication date.