用人脸原型引导情感识别,让模型更懂人类情绪
Grounding Emotion Recognition with Visual Prototypes: VEGA -- Revisiting CLIP in MERC
- 用CLIP图像编码器构建情绪专属人脸锚点,引导特征对齐
- 在IEMOCAP和MELD数据集上达到当前最好性能
- 适合关注多模态情感分析与认知启发模型的研究者
对话中的多模态情感识别因文本、语音和视觉信号的复杂交互而具有挑战性。尽管近期模型通过先进融合策略提升了性能,但往往缺乏心理上合理的先验知识来指导多模态对齐。本文重新审视CLIP,提出一种新的视觉情感引导锚定(VEGA)机制,将类别级视觉语义引入融合与分类过程。不同于以往主要使用CLIP文本编码器的工作,本方法利用其图像编码器,基于面部样本构建情绪特异的视觉锚点。这些锚点引导单模态与多模态特征向具感知基础且心理一致的表示空间对齐,灵感来自认知理论(原型化情绪类别与多感官整合)。采用随机锚点采样策略进一步增强鲁棒性,平衡语义稳定性与类内多样性。在双分支架构中结合自蒸馏,所提出的VEGA增强模型在IEMOCAP和MELD数据集上取得当前最优结果。代码已公开:https://github.com/dkollias/VEGA。
原文摘要 · Abstract (English)
Multimodal Emotion Recognition in Conversations remains a challenging task due to the complex interplay of textual, acoustic and visual signals. While recent models have improved performance via advanced fusion strategies, they often lack psychologically meaningful priors to guide multimodal alignment. In this paper, we revisit the use of CLIP and propose a novel Visual Emotion Guided Anchoring (VEGA) mechanism that introduces class-level visual semantics into the fusion and classification process. Distinct from prior work that primarily utilizes CLIP's textual encoder, our approach leverages its image encoder to construct emotion-specific visual anchors based on facial exemplars. These anchors guide unimodal and multimodal features toward a perceptually grounded and psychologically aligned representation space, drawing inspiration from cognitive theories (prototypical emotion categories and multisensory integration). A stochastic anchor sampling strategy further enhances robustness by balancing semantic stability and intra-class diversity. Integrated into a dual-branch architecture with self-distillation, our VEGA-augmented model achieves sota performance on IEMOCAP and MELD. Code is available at: https://github.com/dkollias/VEGA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。