AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Four Ledgers, Not One Score: Responsible Communication of LLM-Judge Calibration in Biomedical ML

arXiv · AI, language, vision and robotics · article · Sep 14, 2026 · UTC

Synthetic perturbations appear to offer inexpensive calibration data for LLM evaluators in biomedical ML, where expert review is scarce. Yet a planted mutation key is neither a detector output nor automatically human ground truth. We formalize four distinct ledgers: planted perturbations, independent detector outputs, source-linked human dispositions, and human-added discoveries. We then audit the evaluation design, scoring code, read paths, and current human records of a private synthetic Japanese care-handoff workflow. The factory stored 69 planted error cards across 47 targets. Final review

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T11:41:07.830Z. This is not the publication date.