arXiv:2603.13760cs.AIcs.SD2026-03被引 2

通过融合音频与视觉特征,精准预测六类情绪强度,表现优于复杂融合策略。

Multimodal Emotion Regression with Multi-Objective Optimization and VAD-Aware Audio Modeling for the 10th ABAW EMI Track

  • 直接拼接多模态特征,保留各模态特异性。
  • 多目标优化使平均皮尔逊相关系数达0.478567。
  • 适合关注情绪强度建模与多模态融合的研究者。

我们参与了第10届ABAW挑战赛的情感模仿强度(EMI)估计任务,使用Hume-Vidmimic2数据集。该任务旨在预测六种连续情绪维度:钦佩、愉悦、决心、共情痛苦、兴奋和喜悦。通过对预训练高层特征进行系统性多模态探索,我们发现,在当前预训练特征设置下,直接特征拼接优于所测试的复杂融合策略。这一实证发现促使我们设计了一套基于三个核心原则的方法:(i) 通过特征级拼接保留模态特异性;(ii) 通过多目标优化提升训练稳定性与指标一致性;(iii) 采用受VAD启发的潜在先验增强语音表征。最终框架包含基于拼接的多模态融合、共享的六维回归头、结合均方误差、皮尔逊相关系数及辅助分支监督的多目标优化、参数稳定化的EMA机制,以及针对语音分支的VAD-inspired潜先验。在官方验证集上,该方案取得最佳平均皮尔逊相关系数0.478567。

原文摘要 · Abstract (English)

We participated in the 10th ABAW Challenge, focusing on the Emotional Mimicry Intensity (EMI) Estimation track on the Hume-Vidmimic2 dataset. This task aims to predict six continuous emotion dimensions: Admiration, Amusement, Determination, Empathic Pain, Excitement, and Joy. Through systematic multimodal exploration of pretrained high-level features, we found that, under our pretrained feature setting, direct feature concatenation outperformed the more complex fusion strategies we tested. This empirical finding motivated us to design a systematic approach built upon three core principles: (i) preserving modality-specific attributes through feature-level concatenation; (ii) improving training stability and metric alignment via multi-objective optimization; and (iii) enriching acoustic representations with a VAD-inspired latent prior. Our final framework integrates concatenation-based multimodal fusion, a shared six-dimensional regression head, multi-objective optimization with MSE, Pearson-correlation, and auxiliary branch supervision, EMA for parameter stabilization, and a VAD-inspired latent prior for the acoustic branch. On the official validation set, the proposed scheme achieved our best mean Pearson Correlation Coefficient of 0.478567.

多模态情绪识别多目标优化音频建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。