AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

VideoTok4D: A 4D-Aware Video Tokenizer for Compact World Representation

arXiv · AI, language, vision and robotics · article · Sep 11, 2026 · UTC

Video tokenizers have emerged as a cornerstone of modern video modeling, underpinning progress in compression, reconstruction and generation by mapping high-dimensional visual signals into compact latent spaces. However, despite this progress, current tokenization paradigms largely remain within the 2D visual domain, treating videos as image sequences rather than observations of an underlying dynamic 3D world. Consequently, the learned tokens inherit this observation-centric bias, limiting their capacity to compactly represent real-world 4D scenes. To mitigate this issue, we propose VideoTok4D

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T18:22:04.777Z. This is not the publication date.