AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Instruction Distillation: Text Instructions as Visual Examples

arXiv · AI, language, vision and robotics · article · Aug 27, 2026 · UTC

Visual in-context learning (ICL) with multimodal large language models (MLLMs) is effective for fine-grained visual classification, but each retrieved image example consumes several hundred context tokens, making large-$K$ settings prohibitively expensive at inference scale. We propose Instruction Distillation: an offline procedure in which the MLLM itself generates, for each individual training image, a structured identification instruction encoding general appearance cues, features that differentiate the class from visually similar ones, and a common confusion point. Unlike prior work that p

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T08:51:59.673Z. This is not the publication date.