AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Establishing a Spatio-Temporal Language for Scene Representation

CORDIS · observation · Publication date unknown

Establishing a Spatio-Temporal Language for Scene Representation We have recently experienced a boost in AI as the performance of ChatGPT-like large language models has matured from a purely scientific endeavor to deployment in various businesses and real-world applications. Also in computer vision, we have seen tremendous gains that were enabled by scaling to large models trained on vast corpuses of data in an unsupervised fashion. Language is symbolic and can inform about abstract properties and relationships, while vision without human labels does not model explicit semantics and brings distributed representations for spatial structures. Both are complementary, and the fundamental unsolved challenge is to bring them together. The current state of the art is to follow the common paradigm of scale and to naively train models on large amounts of data to exploit the co-occurence of objects in single images and words in text captions to learn their correlation. However, looking at the outputs of these models reveals that they

Read original source ↗ Open in workspace

recordType
award
status
SIGNED
region
EU
value
1563730
unit
EUR

Evidence & attribution

European Commission, CORDIS Horizon Europe project dataset. Metadata adapted.

License: CORDIS reuse policy

First collected: 2026-09-20T04:21:15.460Z. This is not the publication date.