提升多模态情感推理中跨模态感知的可靠性与真实性。
Omni-Perception Policy Optimization for Multimodal Emotion Reasoning

- 用强化学习显式优化多模态感知,分解真实推理为细粒度线索。
- 在多个基准上达到领先性能,显著提升线索利用率与真实性得分。
- 适合关注多模态情感分析、模型可解释性与幻觉抑制的研究者。
我们发现当前面向情感的多模态大模型仍缺乏可靠的跨模态感知:(i) 推理过程中未能充分使用多模态线索,(ii) 表现不忠实,常从其他模态中生成虚假的特定模态陈述。基于此,我们提出 OPPO(Omni-Perception Policy Optimization),一种显式优化多模态感知的强化学习框架。首先,跨模态感知奖励将真实推理分解为细粒度的视觉、声学与情感线索,并奖励能语义恢复这些线索的推理路径。其次,跨模态感知损失对比全模态与单模态屏蔽输入下的策略,仅对特定模态的证据标记施加KL惩罚,以抑制跨模态幻觉。我们进一步构建了 MEP-Bench 诊断基准,用于量化线索利用程度与行为忠实性。实验表明,OPPO 在 MER-UniBench 与 MME-Emotion 上达到最优性能,同时在 MEP-Bench 上显著提升利用度与忠实性分数,凸显充分且忠实的跨模态感知对多模态情感推理的关键作用。
原文摘要 · Abstract (English)
We find that current emotion-oriented Omni-MLLMs still lack reliable omni-modal perception: they (i) underutilize multimodal cues in their reasoning trajectories and (ii) exhibit unfaithful behavior, often hallucinating modality-specific statements from other modalities. Building on these insights, we propose OPPO (Omni-Perception Policy Optimization), a reinforcement learning framework that explicitly optimizes multimodal perception. First, an Omni-Perception Reward decomposes ground-truth reasoning into fine-grained visual, acoustic, and emotion cues and rewards trajectories that semantically recover these cues. Second, an Omni-Perception Loss compares the policy under full and unimodally masked inputs, applying a KL penalty only to modality-specific evidence tokens to suppress cross-modal hallucination. We further introduce MEP-Bench, a diagnostic benchmark that quantifies utilization and faithfulness. Experiments show that OPPO achieves state-of-the-art performance on MER-UniBench and MME-Emotion, while substantially improving utilization and faithfulness scores on MEP-Bench, highlighting the importance of sufficient and faithful omni perception for multimodal emotion reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。