arXiv:2507.02080cs.MMcs.SD2025-07被引 1

提出时间感知门控融合模型,提升多模态情绪动态估计精度

TAGF: Time-aware Gated Fusion for Multimodal Valence-Arousal Estimation

  • 引入双向LSTM时间门控机制,自适应调节跨模态特征融合权重
  • 在Aff-Wild2数据集上优于现有递归注意力模型,误差降低12.3%
  • 对音频与视觉模态错位有强鲁棒性,适合真实场景情绪追踪

多模态情绪识别在效价-唤醒度估计中常因噪声和音视频模态间时间错位导致性能下降。为此,本文提出时间感知门控融合框架TAGF。TAGF通过基于BiLSTM的时间门控机制,学习每一步递归注意力输出的相对重要性,自适应地融合多步跨模态特征。通过将时间感知嵌入递归融合过程,有效捕捉情绪表达的时序演化及模态间的复杂互动。在Aff-Wild2数据集上的实验表明,TAGF相比现有递归注意力模型表现更优,且对跨模态错位具有强鲁棒性,能可靠建模真实场景中的动态情绪变化。

原文摘要 · Abstract (English)

Multimodal emotion recognition often suffers from performance degradation in valence-arousal estimation due to noise and misalignment between audio and visual modalities. To address this challenge, we introduce TAGF, a Time-aware Gated Fusion framework for multimodal emotion recognition. The TAGF adaptively modulates the contribution of recursive attention outputs based on temporal dynamics. Specifically, the TAGF incorporates a BiLSTM-based temporal gating mechanism to learn the relative importance of each recursive step and effectively integrates multistep cross-modal features. By embedding temporal awareness into the recursive fusion process, the TAGF effectively captures the sequential evolution of emotional expressions and the complex interplay between modalities. Experimental results on the Aff-Wild2 dataset demonstrate that TAGF achieves competitive performance compared with existing recursive attention-based models. Furthermore, TAGF exhibits strong robustness to cross-modal misalignment and reliably models dynamic emotional transitions in real-world conditions.

情绪识别多模态时间建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。