AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

SlotDiT: Object-Centric Representations for Diffusion Transformers

arXiv · AI, language, vision and robotics · article · Sep 15, 2026 · UTC

Text-conditioned latent diffusion models perform strongly in video generation and are promising backbones for robotic applications. However, existing approaches rely on pixel-level or VAE-based latent representations that lack explicit semantic structure, leaving the impact of the representation space largely unexplored. Slot-based object-centric representations offer a structured alternative by decomposing scenes into object-level latents, or slots. While they have shown success in dynamics modeling and planning, they have not yet been explored for diffusion-based generative modeling. We intr

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T08:40:59.508Z. This is not the publication date.