AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

SparseTalk - Sparsifying 3D Gaussian Language Fields for Efficient 3D Visual Question Answering

arXiv · AI, language, vision and robotics · article · Sep 14, 2026 · UTC

3D Gaussian language fields provide an explicit, spatially grounded representation for 3D visual question answering (VQA), but their dense semantic features can require tens of thousands of embeddings per scene, resulting in substantial storage, memory, and inference costs. We investigate how much of this representation is actually necessary for downstream reasoning. Starting from a full embedding representation, we systematically sparsify its semantic embeddings, including the previously underexplored regime below a single image-equivalent block down to 8 visual tokens. We compare random, geo

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T11:41:07.830Z. This is not the publication date.