arXiv:2603.11095cs.MMcs.SD2026-03中稿 · ICASSP 2026

解决音视频情绪识别中的帧率不匹配问题,提升跨模态融合效果。

Multimodal Self-Attention Network with Temporal Alignment for Audio-Visual Emotion Recognition

  • 用多模态自注意力编码器同步捕捉音视频内部与跨模态依赖。
  • 在CREMA-D和RAVDESS上优于最新基线,准确率显著提升。
  • 适合关注跨模态时序对齐的音频视频分析研究者。

音频-视觉情绪识别(AVER)方法通常融合话语级特征,即使采用帧级注意力模型也极少处理不同模态间的帧率差异。本文提出一种基于Transformer的框架,聚焦多模态特征的时序对齐。设计采用多模态自注意力编码器,在共享特征空间中同时捕捉模态内与模态间依赖。为应对采样率异构问题,引入时序对齐旋转位置嵌入(TaRoPE),隐式同步音视频标记。此外,提出跨时序匹配(CTM)损失,强制时间邻近样本对的一致性,引导编码器实现更好对齐。在CREMA-D和RAVDESS数据集上的实验表明,该方法持续优于近期基线,说明显式处理帧率不匹配有助于保留时序线索并增强跨模态融合。

原文摘要 · Abstract (English)

Audio-visual emotion recognition (AVER) methods typically fuse utterance-level features, and even frame-level attention models seldom address the frame-rate mismatch across modalities. In this paper, we propose a Transformer-based framework focusing on the temporal alignment of multimodal features. Our design employs a multimodal self-attention encoder that simultaneously captures intra- and inter-modal dependencies within a shared feature space. To address heterogeneous sampling rates, we incorporate Temporally-aligned Rotary Position Embeddings (TaRoPE), which implicitly synchronize audio and video tokens. Furthermore, we introduce a Cross-Temporal Matching (CTM) loss that enforces consistency among temporally proximate pairs, guiding the encoder toward better alignment. Experiments on CREMA-D and RAVDESS datasets demonstrate consistent improvements over recent baselines, suggesting that explicitly addressing frame-rate mismatch helps preserve temporal cues and enhances cross-modal fusion.

情绪识别时序对齐多模态Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。