arXiv:2606.27652cs.AI2026-06被引 1

让快速直觉与慢速推理协同,提升多模态情感识别准确率

MER-R1: Multimodal Emotion Reasoning via Slow-Fast Thinking Synergy

论文配图:MER-R1: Multimodal Emotion Reasoning via Slow-Fast Thinking Synergy
图 1 · 摘自论文原文
  • 通过强化学习框架实现快思与慢思的互补优化
  • 在两个数据集上均达到当前最佳性能,召回率与精确率双提升
  • 适合关注可解释性与高精度情感识别的研究者

我们发现,尽管显式推理能提高多模态情感识别(MER)结果的可解释性,但并不必然提升准确率。具体而言,对于基于推理的多模态大语言模型(MLLMs),直接触发快速回答的快思往往优于经过深思熟虑的慢思。实证分析表明,快思提升召回率,带来更广泛且更自信的预测;而慢思则通过保守过滤错误类别,提高精确率。基于此,我们提出MER-R1,一种将快慢思维互补关系转化为显式优化目标的强化学习框架。双目标解耦分离召回与精确率作为独立优化信号,实现联合优化而非权衡。慢-快信心校准进一步使最终慢思答案与快思直觉对齐,强化正确情感、抑制错误分类。该方法统一了快思的召回导向直觉与慢思的精确导向选择性。我们还提供了理论依据,证明该协同机制可缓解优化过程中的方差干扰。在MER-UniBench与MME-Emotion上的大量实验表明,MER-R1达到当前最优表现,真正实现了推理对情感识别的增益。

原文摘要 · Abstract (English)

We find that explicit reasoning does not necessarily translate into better multimodal emotion recognition (MER) accuracy, even though it makes predictions more interpretable. Specifically, for reasoning-based MLLMs, fast thinking by triggering direct answers often outperforms slow thinking after deliberative reasoning. Our empirical analyses show that fast thinking improves recall with broader and more confident predictions, whereas slow thinking favors precision through conservative filtering of incorrect categories. Building on these insights, we propose MER-R1, a reinforcement learning framework that turns slow-fast complementarity into explicit optimization. Dual-objective disentanglement separates recall and precision into two optimization signals, allowing them to be jointly optimized rather than traded off against each other. Slow-fast confidence calibration further aligns the final slow-thinking answer with fast-thinking intuition, strengthening correct emotions while suppressing incorrect ones. In this way, MER-R1 unifies the recall-oriented intuition of fast thinking with the precision-oriented selectivity of slow thinking. We further provide theoretical justification for this synergy, showing that it mitigates variance-induced interference during optimization. Extensive experiments on MER-UniBench and MME-Emotion show that MER-R1 achieves state-of-the-art performance and makes reasoning genuinely benefit emotion recognition.

多模态情感识别推理协同强化学习可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。