AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

LG-GER: Language-Guided Group Emotion Recognition via Multimodal Evidence Distillation

arXiv · AI, language, vision and robotics · article · Aug 24, 2026 · UTC

Inferring the collective emotional state of a group of people from a single image, a task known as group emotion recognition (GER), requires integrating spatially distributed cues such as faces, poses, interactions, and scene context. Current methods rely on detector-driven multi-stream pipelines. These are trained with only image-level supervision that lacks guidance on which regions matter or how strongly each contributes. We propose LG-GER, a language-guided distillation framework that uses a multimodal large language model (MLLM) to generate dense, spatially grounded evidence, i.e., boundi

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T10:22:00.206Z. This is not the publication date.