AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Integrating Unimodal and Vision-Language Representations in Latent Space for Multi-Label Chest X-Ray Classification

arXiv · AI, language, vision and robotics · article · Aug 31, 2026 · UTC

Multi-label chest X-ray classification has attracted considerable attention in recent years, with the effective use of visual representations and clinical semantic knowledge playing an important role. This study proposes a framework that combines unimodal representations from RAD-DINO with vision--language representations from BioViL-T for the classification of 14 labels in the MIMIC-CXR-JPG dataset. The RAD-DINO and BioViL-T embeddings and their combined representation are refined separately in latent space before being normalized and fused across the three branches. In addition to improving

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T07:22:03.933Z. This is not the publication date.