AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

The record is part of the task: matched-record evaluation of text classifiers across maintenance, safety and recall reporting

arXiv · AI, language, vision and robotics · article · Sep 14, 2026 · UTC

Many operational cases are documented more than once, at different workflow stages and for different purposes, yet model evaluations normally select one of these records before model comparison begins. We treat that selection as part of the evaluation and compare matched records of the same cases under fixed labels and splits in three systems: GE Aerospace repair events, NASA ASRS safety reports and NHTSA vehicle recalls. Across the three GE fields, for events whose label comes from parts transactions independently of the narratives, held-out macro-F1 ranged from 0.33 to 0.91. A difference of

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T09:41:04.278Z. This is not the publication date.