融合视觉、音频与文本,提升真实场景下情绪识别准确率
Multimodal Emotion Recognition via Bi-directional Cross-Attention and Temporal Modeling

- 双向交叉注意力实现视觉与音频特征对称交互
- 引入时间建模与文本对比损失,提升情绪表征质量
- 在真实视频数据上实现0.32的宏平均F1分数
真实场景视频中的表情识别因面部外观、背景条件、音频噪声及情感动态性差异而面临挑战。单一模态(如面部表情或语音)难以捕捉复杂情绪线索。为此,我们提出一种多模态情绪识别框架,用于第10届情感行为分析竞赛(ABAW 10th EXPR任务)。该框架基于大规模预训练模型进行视觉与音频表征学习,并构建统一的多模态架构。为捕捉面部表情序列的时间模式,引入视频窗口内的时序视觉建模。进一步设计双向交叉注意力融合模块,实现视觉与音频特征对称交互,促进跨模态上下文理解与互补信息融合。同时,采用文本引导的对比目标,通过与情绪相关文本提示对齐,引导视觉表示学习语义有意义特征。在ABAW第10届EXPR基准上的实验表明,所提框架有效,宏平均F1得分达0.32,显著优于基线的0.25,验证了时序视觉建模、音频表征学习与跨模态融合在非受限现实环境中的协同优势。
原文摘要 · Abstract (English)
Expression recognition in in-the-wild video data remains challenging due to substantial variations in facial appearance, background conditions, audio noise, and the inherently dynamic nature of human affect. Relying on a single modality, such as facial expressions or speech, is often insufficient for capturing these complex emotional cues. To address this limitation, we propose a multimodal emotion recognition framework for the Expression (EXPR) task in the 10th Affective Behavior Analysis in-the-wild (ABAW) Challenge. Our framework builds on large-scale pre-trained models for visual and audio representation learning and integrates them in a unified multimodal architecture. To better capture temporal patterns in facial expression sequences, we incorporate temporal visual modeling over video windows. We further introduce a bi-directional cross-attention fusion module that enables visual and audio features to interact in a symmetric manner, facilitating cross-modal contextualization and complementary emotion understanding. In addition, we employ a text-guided contrastive objective to encourage semantically meaningful visual representations through alignment with emotion-related text prompts. Experimental results on the ABAW 10th EXPR benchmark demonstrate the effectiveness of the proposed framework, achieving a Macro F1 score of 0.32 compared to the baseline score of 0.25, and highlight the benefit of combining temporal visual modeling, audio representation learning, and cross-modal fusion for robust emotion recognition in unconstrained real-world environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。