AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

MTAC-IFBench: Benchmarking Instruction-Following in Multi-Turn Agentic Coding

arXiv · AI, language, vision and robotics · article · Sep 14, 2026 · UTC

Recently, the rapid development of large language models (LLMs) has reshaped software engineering by enabling autonomous code agents that plan, execute, and utilize external tools iteratively to tackle complex tasks. Beyond achieving functional correctness, these agents must faithfully follow process instructions and constraints throughout the development lifecycle. However, existing benchmarks typically focus on final functional correctness or confine instruction-following evaluation to single-turn, general chat or simple code generation scenarios, leaving instruction-following in multi-turn

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T11:41:07.830Z. This is not the publication date.