arXiv:2503.10603cs.CVcs.AI2025-03被引 2

提出双阶段跨模态对齐框架,提升复杂环境下的情感模仿强度估计精度。

Technical Approach for the EMI Challenge in the 8th Affective Behavior Analysis in-the-Wild Competition

  • 分两阶段对齐视觉、音频与文本特征,增强跨模态协同。
  • 在Hume-Vidmimic2数据集上验证集平均相关系数达0.51,测试集达0.68。
  • 适合关注细粒度情感分析与多模态融合的研究者。

情感模仿强度(EMI)估计在理解人类社交行为和推动人机交互中具有关键作用。核心挑战在于动态相关性建模与多模态时序信号的鲁棒融合。针对现有方法在跨模态协同利用不足、对噪声敏感及细粒度对齐能力受限的问题,本文提出一种双阶段跨模态对齐框架。第一阶段基于CLIP架构构建视觉-文本与音频-文本对比学习网络,通过模态解耦预训练实现初步特征空间对齐。第二阶段引入时序感知动态融合模块,结合时间卷积网络(TCN)捕捉面部表情的宏观演变模式,以及门控双向LSTM处理声学特征的局部动态。此外,提出一种质量引导融合策略,在遮挡和噪声条件下实现可微权重分配以补偿模态差异。在Hume-Vidmimic2数据集上的实验表明,该方法在验证集上六种情绪维度的平均皮尔逊相关系数为0.51,测试集达到0.68,位列第8届ABAW竞赛的EMI挑战赛道亚军,为开放环境下细粒度情感分析提供了新路径。

原文摘要 · Abstract (English)

Emotional Mimicry Intensity (EMI) estimation plays a pivotal role in understanding human social behavior and advancing human-computer interaction. The core challenges lie in dynamic correlation modeling and robust fusion of multimodal temporal signals. To address the limitations of existing methods--insufficient exploitation of cross-modal synergies, sensitivity to noise, and constrained fine-grained alignment capabilities--this paper proposes a dual-stage cross-modal alignment framework. Stage 1 develops vision-text and audio-text contrastive learning networks based on a CLIP architecture, achieving preliminary feature-space alignment through modality-decoupled pre-training. Stage 2 introduces a temporal-aware dynamic fusion module integrating Temporal Convolutional Networks (TCN) and gated bidirectional LSTM to capture macro-evolution patterns of facial expressions and local dynamics of acoustic features, respectively. A novel quality-guided fusion strategy further enables differentiable weight allocation for modality compensation under occlusion and noise. Experiments on the Hume-Vidmimic2 dataset demonstrate superior performance with an average Pearson correlation coefficient of 0.51 across six emotion dimensions on the validate set. Remarkably, our method achieved 0.68 on the test set, securing runner-up in the EMI Challenge Track of the 8th ABAW (Affective Behavior Analysis in the Wild) Competition, offering a novel pathway for fine-grained emotion analysis in open environments.

情感识别多模态融合跨模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。