融合视觉、行为与音频,提升真实场景下情绪连续识别准确率。
Team RAS in 10th ABAW Competition: Multimodal Valence and Arousal Estimation Approach
- 用多模态融合建模面部、行为与音频特征,增强鲁棒性。
- 在Aff-Wild2数据集上达到0.658的CCC得分,优于基准方法。
- 适合研究真实场景下情绪识别或跨模态融合的开发者参考。
在真实场景中进行持续情绪识别(以愉悦度和唤醒度为指标)仍具挑战,因外观、头部姿态、光照、遮挡及个体表达差异大。本文提出一种多模态方法,融合面部、行为与音频信息。面部模态基于GRADA帧级嵌入与Transformer时序回归;使用Qwen3-VL-4B-Instruct从视频片段提取行为信息,Mamba模型捕捉跨片段时序动态;音频模态采用WavLM-Large结合注意力统计池化,并引入跨模态过滤阶段以剔除不可靠或非语音段。通过两种融合策略:有向交叉模态专家混合(适应性加权交互)与可靠性感知音视频融合(帧级视觉+音频补充)。实验在遵循第10届ABAW挑战协议的Aff-Wild2数据集上进行,结果表明所提融合策略在开发集上实现0.658的一致性相关系数(CCC)。
原文摘要 · Abstract (English)
Continuous emotion recognition in terms of valence and arousal under in-the-wild (ITW) conditions remains a challenging problem due to large variations in appearance, head pose, illumination, occlusions, and subject-specific patterns of affective expression. We present a multimodal method for valence-arousal estimation ITW. Our method combines three complementary modalities: face, behavior, and audio. The face modality relies on GRADA-based frame-level embeddings and Transformer-based temporal regression. We use Qwen3-VL-4B-Instruct to extract behavior-relevant information from video segments, while Mamba is used to model temporal dynamics across segments. The audio modality relies on WavLM-Large with attention-statistics pooling and includes a cross-modal filtering stage to reduce the influence of unreliable or non-speech segments. To fuse modalities, we explore two fusion strategies: a Directed Cross-Modal Mixture-of-Experts Fusion Strategy that learns interactions between modalities with adaptive weighting, and a Reliability-Aware Audio-Visual Fusion Strategy that combines visual features at the frame-level while using audio as complementary context. The results are reported on the Aff-Wild2 dataset following the 10th Affective Behavior Analysis in-the-Wild (ABAW) challenge protocol. Experiments demonstrate that the proposed multimodal fusion strategy achieves a Concordance Correlation Coefficient (CCC) of 0.658 on the Aff-Wild2 development set.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。