AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

When Is Noise Response Universal? Tokenization as the Hidden Variable in Language Models

arXiv · AI, language, vision and robotics · article · Aug 26, 2026 · UTC

The performance of textual neural models often degrades when their inputs are corrupted by noise such as typos, OCR errors, or dropped words. We study the degradation rate across neural models, both sentence embeddings and decoder-only LLMs, and find that how consistent it is depends on the scale of the noise: under word-level noise, models with very different architectures decline along nearly the same curve, while under character-level noise they separate. We further identify the determining factor to be the training objective, not the architecture: eight encoders spanning six pretraining pa

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T09:11:58.312Z. This is not the publication date.