SOURCE-LINKED INTELLIGENCE
Unified Transcription and Translation for Extended Reality
Unified Transcription and Translation for Extended Reality The aim of UTTER is to leverage large language models to build the next generation of multimodal eXtended reality (XR) technologies for transcription, translation, summarisation, and minuting. We will make these technologies scalable, adaptable, contextualised, robust, explainable, and emotion-aware. We will increase the context-sensitivity of the technologies, so they can take into account the full history of the conversation, as well as its wider context. We will introduce confidence-aware models, which can take into account their own limitations. We will develop explainable models, so the human user can know why the model made the decisions it did. We will improve adaptation, so that domain-specific and language-specific models can be quickly rolled out. For these advances we will make use of pre-trained eXtended reality (XR) models, which optimally combine text and speech signals, and are trained efficiently with
Read original source ↗ Open in workspace
- recordType
- award
- status
- SIGNED
- region
- EU
- value
- 4070321.89
- unit
- EUR
Evidence & attribution
European Commission, CORDIS Horizon Europe project dataset. Metadata adapted.
License: CORDIS reuse policy
First collected: 2026-09-20T00:21:03.701Z. This is not the publication date.