AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks

arXiv · AI, language, vision and robotics · article · Sep 15, 2026 · UTC

Benchmarks are central to how progress in large language models (LLMs) is assessed and communicated. Yet model rankings alone reveal little about how evaluation requirements themselves are changing. The expanding variety of benchmarks offers another perspective: what researchers expect LLMs to do, and what they count as successful performance. We systematically map 14,767 papers introducing or updating evaluation resources from arXiv submissions between January 2022 and August 2026. Using staged screening and automated full-text coding, we examine changes in target systems and domains, evaluat

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T08:40:59.508Z. This is not the publication date.