SOURCE-LINKED INTELLIGENCE
Reproducibility is not construct validity: LLM measurement of institutionally situated communication
High annotation reproducibility does not necessarily imply that an LLM-inferred measure captures the construct it is intended to measure. We test this distinction using a dataset from the European Commission's AI Act consultation, linking structured survey responses to free-text consultation submissions from the same stakeholders. LLM annotations of consultation submissions are highly reproducible (intraclass correlations > 0.99), yet show limited convergence with survey-reported measures of the nominal construct they were intended to approximate. Divergence between survey-and LLM-inferred tex
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-17T08:19:58.000Z
- arXiv · Artificial Intelligence · 2026-09-17T08:19:58.000Z
First collected: 2026-09-19T20:26:32.566Z. This is not the publication date.