用语言模型生成情感证据,让图像识群情更准更快。
LG-GER: Language-Guided Group Emotion Recognition via Multimodal Evidence Distillation

- 用大语言模型生成带位置的情感证据,指导模型学习
- 在两个数据集上超越检测依赖的现有方法
- 推理无需检测器,适合实时和低资源场景
从单张图像推断群体情绪状态(即群情识别,GER)需融合面部、姿态、互动和场景等空间分布线索。现有方法依赖检测器驱动的多流流水线,仅使用图像级监督,缺乏对关键区域及其贡献强度的指引。本文提出 LG-GER,一种基于语言引导的多模态证据蒸馏框架,利用多模态大语言模型(MLLM)为训练图像生成密集、空间定位的证据——即带有情绪信号与置信度的边界框。这些结构化证据通过四类互补损失(分类、区域-文本对齐、空间情绪、空间置信度回归)蒸馏至单一视觉-语言模型(VLM)主干网络。推理阶段无需检测器、无需MLLM、无需多流融合,使GER可部署于实时与资源受限场景。LG-GER在两个基准数据集(GroupEmoW 和 GAF~3.0)上评估,性能达到或优于依赖检测与多流处理的先进方法。
原文摘要 · Abstract (English)
Inferring the collective emotional state of a group of people from a single image, a task known as group emotion recognition (GER), requires integrating spatially distributed cues such as faces, poses, interactions, and scene context. Current methods rely on detector-driven multi-stream pipelines. These are trained with only image-level supervision that lacks guidance on which regions matter or how strongly each contributes. We propose LG-GER, a language-guided distillation framework that uses a multimodal large language model (MLLM) to generate dense, spatially grounded evidence, i.e., bounding boxes paired with emotion signals and confidence scores, for the training images. This structured evidence is distilled into a single vision-language model (VLM) backbone through four complementary losses: classification, region-text grounding, spatial emotion, and spatial confidence regression. At inference, LG-GER requires no detectors, no MLLM, and no multi-stream fusion, making GER practical for real-time and resource-constrained deployment. LG-GER has been evaluated on two benchmark GER datasets (GroupEmoW and GAF~3.0) and achieves competitive or superior results compared to state-of-the-art methods that require detection and multi-stream processing at inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。