用文本为中心的多模态融合方法,提升视频级犹豫与矛盾情绪识别准确率。
Team RAS in 11th ABAW Competition: Multimodal Ambivalence Recognition Approach

- 以文本为锚点,通过门控残差调整融合语音、表情和场景特征
- 在私有测试集上达到78.24%的平均宏F1,优于纯文本模型4.03%
- 无需复杂模型集成,适合实际部署的轻量级情绪识别系统
自动识别犹豫与矛盾情绪具有挑战性,因这些状态可能表现为语言、语音、面部和情境模式不一致。现有高性能系统常依赖计算开销大的集成模型。本文提出一种单模型文本中心的多模态方法,用于第11届在野情感与行为分析(ABAW)挑战赛中的视频级犹豫与矛盾情绪识别。该方法通过文本中心多模态融合模型结合语言、语音、面部和场景特征。文本残差融合(Text Residual Fusion)以文本为锚点,基于其他模态进行门控残差调整。在行为矛盾/犹豫(BAH)语料库上的实验表明,文本是表现最强的单模态。该模型在开发集与公开测试集上平均宏F1得分为75.14%,在私有测试集上达78.24%,比纯文本模型高出4.03%。结果表明,互补的多模态信息可有效提升识别性能,而无需大规模模型集成。
原文摘要 · Abstract (English)
Automatic recognition of ambivalence and hesitancy is challenging because these states may be expressed through inconsistent linguistic, acoustic, facial, and contextual patterns, while top-performing systems often rely on computationally expensive ensembles. We present a single text-centered multimodal approach for video-level ambivalence and hesitancy recognition for the 11th Affective & Behavior Analysis in-the-Wild (ABAW) Challenge. The proposed approach combines linguistic, acoustic, facial, and scene features using text-centered multimodal fusion model. Text Residual Fusion treats text as the anchor modality and applies gated residual adjustments based on the other modalities. Experiments on the Behavioural Ambivalence/Hesitancy (BAH) corpus confirm that text is the strongest unimodal modality. The Text Residual Fusion model achieves an average Macro F1-score (MF1) of 75.14% across the Development and Public Test subsets. On the Private Test subset, it reaches an MF1 of 78.24%, outperforming the text model by 4.03%. These results demonstrate that complementary multimodal information can improve recognition performance without requiring a large model ensemble.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。