SOURCE-LINKED INTELLIGENCE
DeltaSelect: Affordable A/B Testing for Coding Agents
Coding-agent benchmarks are built for broad and comprehensive comparisons, not frequent development decisions. Individual runs vary, full suites are expensive, and the benchmark harness may differ from the harness used in practice. In a resampling analysis of DeepSWE's published trials, only 19.5% of tasks (22 of 113) had a fifth-percentile Pearson correlation of at least 0.50 with full-benchmark performance. The paper presents DeltaSelect, an open-source method that identifies tasks whose one-run results consistently track full-benchmark performance using Pearson correlation, maps fractional
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · Artificial Intelligence · 2026-09-17T02:42:02.000Z
- arXiv · AI, language, vision and robotics · 2026-09-17T02:42:02.000Z
First collected: 2026-09-19T20:26:32.566Z. This is not the publication date.