arXiv:2503.12261cs.CVcs.SD2025-03被引 2

提出自适应门控机制,提升音视频情绪识别在弱互补情况下的表现。

United we stand, Divided we fall: Handling Weak Complementary Relationships for Audio-Visual Emotion Recognition in Valence-Arousal Space

  • 设计门控递归联合交叉注意力,动态选择强弱互补特征
  • 在Affwild2数据集上,情绪识别的CCC达0.561(测试集)
  • 适合处理音视频模态互补性不强的真实场景应用

音频和视觉模态是视频中两种主要的非接触式通道,通常预期具有互补关系。然而,二者未必总能互补,导致音视频特征表示效果不佳。本文提出门控递归联合交叉注意力(GRJCA),通过门控机制自适应选择最相关特征,有效捕捉音视频模态间的协同关系。具体地,在递归联合交叉注意力(RJCA)基础上引入门控机制,根据模态间互补强度控制多轮迭代中输入特征与注意力特征的信息流动:若互补性强,则强调交叉注意力特征;否则保留原始特征。为进一步提升性能,还采用分层门控策略,在每轮迭代后及输出层面均引入门控。所提方法增强了对弱互补关系的处理灵活性。在挑战性的Affwild2数据集上进行大量实验,结果表明,该模型在测试集上实现情感维度(效价与唤醒度)的协方差相关系数(CCC)分别为0.561和0.620,在验证集上分别为0.623和0.660。

原文摘要 · Abstract (English)

Audio and visual modalities are two predominant contact-free channels in videos, which are often expected to carry a complementary relationship with each other. However, they may not always complement each other, resulting in poor audio-visual feature representations. In this paper, we introduce Gated Recursive Joint Cross Attention (GRJCA) using a gating mechanism that can adaptively choose the most relevant features to effectively capture the synergic relationships across audio and visual modalities. Specifically, we improve the performance of Recursive Joint Cross-Attention (RJCA) by introducing a gating mechanism to control the flow of information between the input features and the attended features of multiple iterations depending on the strength of their complementary relationship. For instance, if the modalities exhibit strong complementary relationships, the gating mechanism emphasizes cross-attended features, otherwise non-attended features. To further improve the performance of the system, we also explored a hierarchical gating approach by introducing a gating mechanism at every iteration, followed by high-level gating across the gated outputs of each iteration. The proposed approach improves the performance of RJCA model by adding more flexibility to deal with weak complementary relationships across audio and visual modalities. Extensive experiments are conducted on the challenging Affwild2 dataset to demonstrate the robustness of the proposed approach. By effectively handling the weak complementary relationships across the audio and visual modalities, the proposed model achieves a Concordance Correlation Coefficient (CCC) of 0.561 (0.623) and 0.620 (0.660) for valence and arousal respectively on the test set (validation set).

音视频识别情绪分析多模态融合门控机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。