多模态融合预测六种情绪强度,效果优于多数基线。
Two-Stage Multimodal Framework for Emotion Mimicry Intensity Prediction
- 分两阶段训练,先独立编码文本、音频、视觉特征,再轻量融合。
- 最佳模型在验证集上平均皮尔逊相关系数达0.4722,测试集达0.57。
- 支持可复现基线,适合情绪分析与多模态建模研究者参考。
我们提交了Hume-ABAW10情感模仿强度(EMI)挑战赛的方案,目标是从真实场景的多模态视频片段中预测六种连续情绪强度:钦佩、愉悦、决心、共情痛苦、兴奋和喜悦。提出一种分阶段多模态框架,融合文本、音频与视觉表示,并可选加入运动分支。方法先独立训练各模态编码器,再通过带模态丢弃和受控编码器适配的轻量回归器进行融合。在所有提交系统中,文本-音频-视觉-运动融合模型在扩展4:1划分下表现最佳,平均皮尔逊相关系数为0.4722。尽管运动分支仅带来微小提升,但其行为值得深入研究。团队在挑战赛中位列第三,测试集平均皮尔逊相关系数为0.57。整体提供了一个实用且可复现的EMI预测基线。
原文摘要 · Abstract (English)
We present our submission to the Hume-ABAW10 Emotional Mimicry Intensity (EMI) Challenge, which aims to predict six continuous emotion intensity dimensions: Admiration, Amusement, Determination, Empathic Pain, Excitement, and Joy, from in-the-wild multimodal video clips. We propose a staged multimodal framework that combines textual, acoustic, and visual representations, with an optional motion branch. Our approach first trains modality-specific encoders independently and then fuses their learned representations through a lightweight regressor with modality dropout and controlled encoder adaptation. Across our submitted systems, the best validation performance is obtained by the text--audio--vision--motion fusion model under the expanded 4:1 split, achieving an average Pearson correlation of 0.4722. Although the motion branch yields only very slight gains, its behavior can be interesting to study. Our team was placed third in the EMI challenge, achieving an average Pearson correlation of 0.57 for the test set. Overall, we provide a practical and reproducible baseline for EMI prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。