arXiv:2503.12623cs.LGcs.AI2025-03CVPR被引 14

多模态注意力模型提升真实场景情绪识别准确率

MAVEN: Multi-modal Attention for Valence-Arousal Emotion Network

  • 用双向跨模态注意力融合视觉、音频、文本信息
  • 在Aff-Wild2数据集上情感相关性系数达0.3061
  • 适合研究真实对话中情绪动态变化的学者

野外动态情绪识别因情绪表达短暂及多模态线索时间错位而具挑战性。传统方法常忽略愉悦度与唤醒度之间的内在关联。本文提出的多模态注意力情绪网络(MAVEN)通过双向跨模态注意力机制整合视觉、音频和文本模态。MAVEN使用模态专用编码器从同步视频帧、音频片段和转录文本中提取特征,基于Russell环形模型以极坐标形式预测情绪。在Aff-Wild2数据集上的评估显示,其一致性相关系数(CCC)达0.3061,优于ResNet-50基线模型的0.22。多阶段架构捕捉对话视频中细微且瞬时的情绪变化,提升真实场景下的情绪识别性能。代码已开源。

原文摘要 · Abstract (English)

Dynamic emotion recognition in the wild remains challenging due to the transient nature of emotional expressions and temporal misalignment of multi-modal cues. Traditional approaches predict valence and arousal and often overlook the inherent correlation between these two dimensions. The proposed Multi-modal Attention for Valence-Arousal Emotion Network (MAVEN) integrates visual, audio, and textual modalities through a bi-directional cross-modal attention mechanism. MAVEN uses modality-specific encoders to extract features from synchronized video frames, audio segments, and transcripts, predicting emotions in polar coordinates following Russell's circumplex model. The evaluation of the Aff-Wild2 dataset using MAVEN achieved a concordance correlation coefficient (CCC) of 0.3061, surpassing the ResNet-50 baseline model with a CCC of 0.22. The multistage architecture captures the subtle and transient nature of emotional expressions in conversational videos and improves emotion recognition in real-world situations. The code is available at: https://github.com/Vrushank-Ahire/MAVEN_8th_ABAW

情绪识别多模态注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。