AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Certifying Model Upgrades with Slice-Wise Non-Regression and Incumbent Fallback

arXiv · AI, language, vision and robotics · article · Sep 12, 2026 · UTC

An updated model can improve an aggregate metric while degrading a slice that matters to a downstream user. We study checkpoint selection subject to non-regression tolerances relative to a retained incumbent. The central distinction is between failing to detect harm and certifying non-inferiority: the former can release harmful updates with high probability when evaluation is noisy. We give a reproducible release procedure that separates candidate search from independent, paired evaluation and returns the exact incumbent when certification fails. Applying established intersection-union and Lea

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T16:41:15.630Z. This is not the publication date.