AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

ProtoLIP: From Sentence-Level to Object-Level Evidence Disentanglement

arXiv · AI, language, vision and robotics · article · Sep 14, 2026 · UTC

Query-conditioned vision--language models enable fine-grained interpretation by revealing how visual evidence changes with textual queries. However, evidence conditioned on complete descriptions does not necessarily resolve into object-specific evidence, nor does an exposed evidence map necessarily identify the evidence that constitutes the model's prediction. Across multiple VLM architectures and independent benchmarks, we find that object-level queries often retain evidence from co-occurring objects and shared context. In this paper, we introduce \textbf{ProtoLIP}, a lightweight prototype-me

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T09:41:04.278Z. This is not the publication date.