arXiv:2603.23650cs.CV2026-03

融合12个模型,用多模态数据识别混合情绪并预测情感强度。

Foundation Model Embeddings Meet Blended Emotions: A Multimodal Fusion Approach for the BLEMORE Challenge

  • 用六类编码器晚期概率融合,引入Gemini视频嵌入提升表现。
  • 仅2秒输入即达0.320准确率,关键音频层选6-12效果最优。
  • 个性化表达差异大是主要瓶颈,适合做情绪分析的工程应用。

我们在FG 2026的BLEMORE挑战中提出一种混合情绪识别与相对显著性预测系统。方法结合六个编码器族进行后期概率融合:基于S4D-ViTMoE的脸部编码器经软标签KL训练优化,冻结层选择的Wav2Vec2音频特征,微调的肢体语言编码器(TimeSformer、VideoMAE),以及首次用于情绪识别的Gemini Embedding 2.0——其视频嵌入仅需2秒输入即可实现0.320的高存在准确率(ACCP)。实验发现:从冻结的Wav2Vec2中选择6-12层音调编码优于端到端微调(得分0.207 vs. 0.161),因非语言音频在该任务中语音层不重要;后处理显著性阈值β在不同折间波动于0.05至0.43,揭示个性化表达风格是核心瓶颈;任务适配编码器在集成中占总权重62%。12编码器系统在测试集上得分为0.279(ACCP=0.391,ACCS=0.168),位列第6。

原文摘要 · Abstract (English)

We present our system for the BLEMORE Challenge at FG 2026 on blended emotion recognition with relative salience prediction. Our approach combines six encoder families through late probability fusion: an S4D-ViTMoE face encoder adapted with soft-label KL training, frozen layer-selective Wav2Vec2 audio features, finetuned body-language encoders (TimeSformer, VideoMAE), and -- for the first time in emotion recognition -- Gemini Embedding 2.0, a large multimodal model whose video embeddings produce competitive presence accuracy (ACCP = 0.320) from only 2 seconds of input. Three key findings emerge from our experiments: selecting prosody-encoding layers (6--12) from frozen Wav2Vec2 outperforms end-to-end finetuning (Score 0.207 vs. 0.161), as the non-verbal nature of BLEMORE audio makes phonetic layers irrelevant; the post-processing salience threshold $β$ varies from 0.05 to 0.43 across folds, revealing that personalized expression styles are the primary bottleneck; and task-adapted encoders collectively receive 62\% of ensemble weight over general-purpose baselines. Our 12-encoder system achieves Score = 0.279 (ACCP = 0.391, ACCS = 0.168) on the test set, placing 6th.

情绪识别多模态融合Gemini音频特征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。