arXiv:2607.14702cs.CVcs.AI2026-07被引 2

用文本为中心的多模态融合方法,提升视频级犹豫与矛盾情绪识别准确率。

Team RAS in 11th ABAW Competition: Multimodal Ambivalence Recognition Approach

论文配图:Team RAS in 11th ABAW Competition: Multimodal Ambivalence Recognition Approach
图 1 · 摘自论文原文
  • 以文本为锚点,通过门控残差调整融合语音、表情和场景特征
  • 在私有测试集上达到78.24%的平均宏F1,优于纯文本模型4.03%
  • 无需复杂模型集成,适合实际部署的轻量级情绪识别系统

自动识别犹豫与矛盾情绪具有挑战性,因这些状态可能表现为语言、语音、面部和情境模式不一致。现有高性能系统常依赖计算开销大的集成模型。本文提出一种单模型文本中心的多模态方法,用于第11届在野情感与行为分析(ABAW)挑战赛中的视频级犹豫与矛盾情绪识别。该方法通过文本中心多模态融合模型结合语言、语音、面部和场景特征。文本残差融合(Text Residual Fusion)以文本为锚点,基于其他模态进行门控残差调整。在行为矛盾/犹豫(BAH)语料库上的实验表明,文本是表现最强的单模态。该模型在开发集与公开测试集上平均宏F1得分为75.14%,在私有测试集上达78.24%,比纯文本模型高出4.03%。结果表明,互补的多模态信息可有效提升识别性能,而无需大规模模型集成。

原文摘要 · Abstract (English)

Automatic recognition of ambivalence and hesitancy is challenging because these states may be expressed through inconsistent linguistic, acoustic, facial, and contextual patterns, while top-performing systems often rely on computationally expensive ensembles. We present a single text-centered multimodal approach for video-level ambivalence and hesitancy recognition for the 11th Affective & Behavior Analysis in-the-Wild (ABAW) Challenge. The proposed approach combines linguistic, acoustic, facial, and scene features using text-centered multimodal fusion model. Text Residual Fusion treats text as the anchor modality and applies gated residual adjustments based on the other modalities. Experiments on the Behavioural Ambivalence/Hesitancy (BAH) corpus confirm that text is the strongest unimodal modality. The Text Residual Fusion model achieves an average Macro F1-score (MF1) of 75.14% across the Development and Public Test subsets. On the Private Test subset, it reaches an MF1 of 78.24%, outperforming the text model by 4.03%. These results demonstrate that complementary multimodal information can improve recognition performance without requiring a large model ensemble.

情绪识别多模态融合轻量化模型文本中心

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。