AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Bigger Text Encoders Can Hurt CLIP Zero-Shot Performance

arXiv · AI, language, vision and robotics · article · Sep 4, 2026 · UTC

Contrastive Language-Image Pretraining (CLIP) is a building block of many machine learning applications. Scaling laws have guided resource allocation for large-scale training, yet prior work treats total CLIP model size as a single variable, without exploring how the capacity split between encoders impacts downstream performance. Here, we train multiple CLIP models with different vision and text encoder sizes, revealing that for most vision encoders, there is an optimal text encoder size beyond which zero-shot performance degrades---even as total parameter count increases. Exploiting this beha

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T21:52:07.471Z. This is not the publication date.