AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Skill-based Agentic Evaluation for Real-time Data Science Tasks

arXiv · AI, language, vision and robotics · article · Sep 15, 2026 · UTC

We present a framework for evaluating data-science agents on live, continuously updated data using executable ground truth and format-agnostic factoid scoring. Consider this example query: "what were last week's audience sizes"---the reference answer changes as the underlying data changes, so static references become outdated and standard LLM-as-a-judge pipelines cannot verify responses against a fixed ground truth. Our central contribution, ground-truth-as-code, encodes each expected answer as an executable reference function that recomputes the answer directly from live data at evaluation ti

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T09:01:24.920Z. This is not the publication date.