SOURCE-LINKED INTELLIGENCE
Same Patient, Different Order: Action-Level Reliability of Clinical LLM Agents Under Repeated Runs
A clinical agent benchmark can report the same verdict on identical inputs while the agent files a materially different order on each run. Such agents order tests, request medications and place referrals, yet benchmarks typically score one run per task and rarely ask whether identical inputs produce identical actions; MedAgentBench, the benchmark we use, scores a single attempt and says so. To measure this gap we introduce "same-input rerun", which replays a task with every input held fixed and compares the orders rather than the score, with six reliability metrics, and apply it to 1000 MedAge
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-11T22:44:35.000Z
First collected: 2026-09-20T16:41:15.630Z. This is not the publication date.