AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

BenchMIRT: What are LLM benchmarks actually measuring?

Ai2 Research · article · Sep 1, 2026 · UTC

BenchMIRT is a new method for auditing LLM benchmarks question by question, revealing which capabilities they actually measure and helping researchers build smaller, more focused, and easier-to-interpret evaluations.

Read original source ↗ Open in workspace

recordType
article
region
Global

Evidence & attribution

First collected: 2026-09-19T20:26:46.936Z. This is not the publication date.

Observed changes

AIIC observation times, not verified publisher revision times. Up to eight recent revisions.

2026-09-20T20:32:20.942Z

  • publishedAt: 2026-09-01T07:00:00.000Z2026-09-01T00:00:00.000Z