AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Don't Just Look, Intervene: Perturbation Based Region Labeling for VQA Images

arXiv · AI, language, vision and robotics · article · Sep 1, 2026 · UTC

Vision Language Models (VLMs) should rely on visual evidence that directly determines the correct answer, but supervision for grounding visual reasoning is often expensive to obtain manually or tied to dataset-specific annotation primitives. We instead introduce model-causal visual evidence as an annotation target, defined as the set of image regions whose counterfactual intervention changes a model's answer distribution for a given image-question pair. Based on this principle, we introduce Counterfactual Search for Grounding Regions (CSGR). CSGR is a scalable pipeline that proposes candidate

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T05:51:54.566Z. This is not the publication date.