SOURCE-LINKED INTELLIGENCE
Graphs without Labels: Multimodal Structure Learning without Human Supervision
raining models with data in more than one modality, such as videos capturing visual and audio information or documents containing image and text. Current approaches use such data to train large-scale deep learning models without human supervision by sampling pair-wise data e.g., an image-text pair from a website and train the network e.g. to identify matching vs. not matching pairs to learn better representations. We argue that multimodal learning can do more: by combining information from different sources, multimodal models capture cross-modal semantic entities, and as most multimodal documents are a collection of connected modalities and topics, multimodal models should allow us to capture the inherent high-level topology of such data. The goal of the following project is to learn semantic structures from multimodal data to capture long-range concepts and relations in multimodal data via multimodal and self-supervision learning without human annotation. We will represent this information in form of a graph, considering latent semantic concepts as nodes and their connectivity as ed
Read original source ↗ Open in workspace
- recordType
- award
- status
- SIGNED
- region
- EU
- value
- 1499438
- unit
- EUR
Evidence & attribution
European Commission, CORDIS Horizon Europe project dataset. Metadata adapted.
License: CORDIS reuse policy
First collected: 2026-09-20T02:21:08.944Z. This is not the publication date.