AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention

arXiv · AI, language, vision and robotics · article · Sep 14, 2026 · UTC

Diffusion Transformers deliver state-of-the-art video generation, but their long spatiotemporal sequences make attention the dominant deployment cost, and a deployable low-bit kernel must be accurate and fast. Accuracy is limited by outliers: a block's quantization scale is set by its largest entries, leaving typical entries confined to a narrow range of representable values. Prior work smooths queries and keys, but value outliers follow no fixed channel or spatiotemporal structure and remain the dominant source of output error. Speed is limited by softmax: low-bit Tensor Cores accelerate only

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T09:41:04.278Z. This is not the publication date.