AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

SynthSentry: Detecting Synthetic Data Contamination in Language Model Training Data

arXiv · AI, language, vision and robotics · article · Sep 11, 2026 · UTC

Large language models trained recursively on their own or other models' outputs undergo model collapse, in which distributional tails and factual accuracy deteriorate while fluency survives. Prior work diagnoses collapse after training; the actionable problem is screening a corpus of unknown provenance before training. We introduce SynthSentry, a corpus-level, model-agnostic contamination signal requiring no access to the generating model, no generation history, and no synthetic labels. The score is a distributional divergence over three statistics: lexical diversity collapse, n-gram tail trun

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T18:42:18.733Z. This is not the publication date.