AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

When Modality Gap Reduction Fails: Prediction-Level Hubness in CLIP

arXiv · AI, language, vision and robotics · article · Sep 1, 2026 · UTC

Reducing the modality gap between image and text representations in CLIP is widely expected to improve cross-modal alignment and downstream performance. However, a smaller average image-text gap does not necessarily lead to consistent accuracy gains. We analyze this mismatch from the perspective of the decision structure in zero-shot classification, i.e. selecting the most similar class-text prototype for an input image. Zero-shot accuracy depends not only on average image--text alignment, but also on class-wise decision margins. Using Linear correction as an analytically tractable case, we sh

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T06:11:57.537Z. This is not the publication date.