SOURCE-LINKED INTELLIGENCE
Establishing a Spatio-Temporal Language for Scene Representation
Establishing a Spatio-Temporal Language for Scene Representation We have recently experienced a boost in AI as the performance of ChatGPT-like large language models has matured from a purely scientific endeavor to deployment in various businesses and real-world applications. Also in computer vision, we have seen tremendous gains that were enabled by scaling to large models trained on vast corpuses of data in an unsupervised fashion. Language is symbolic and can inform about abstract properties and relationships, while vision without human labels does not model explicit semantics and brings distributed representations for spatial structures. Both are complementary, and the fundamental unsolved challenge is to bring them together. The current state of the art is to follow the common paradigm of scale and to naively train models on large amounts of data to exploit the co-occurence of objects in single images and words in text captions to learn their correlation. However, looking at the outputs of these models reveals that they
Read original source ↗ Open in workspace
- recordType
- award
- status
- SIGNED
- region
- EU
- value
- 1563730
- unit
- EUR
Evidence & attribution
European Commission, CORDIS Horizon Europe project dataset. Metadata adapted.
License: CORDIS reuse policy
First collected: 2026-09-20T04:21:15.460Z. This is not the publication date.