用心理理论提升多模态模型的情感理解能力
Unveiling the Cognitive Compass: Theory-of-Mind-Guided Multimodal Emotion Reasoning
- 构建分层基准测试,诊断模型在认知深度上的能力断点
- 引入心理状态追踪机制,显著提升情感推理准确率与合理性
- 适合研究情感智能、认知推理的AI开发者和研究员
尽管多模态大模型进展迅速,其深层情感理解能力仍受限。我们认为,真正的情感能力需显式建模心理理论(ToM),即情绪产生的认知基础。为此,我们提出HitEmotion——一个基于心理理论的分层基准,用于诊断模型在不同认知深度下的能力断点。同时,设计了一种心理理论引导的推理链,通过追踪心理状态并校准跨模态证据,实现更真实的感情推理。此外,提出TMPO强化学习方法,以中间心理状态作为过程监督信号,指导并增强模型推理。大量实验表明,HitEmotion揭示了当前先进模型在高阶认知任务中的深层情感推理缺陷。在评估中,心理理论引导的推理链与TMPO均提升了最终任务准确率,并生成更可信、更连贯的推理过程。本工作为社区提供了评估与增强多模态大模型认知性情感能力的实用工具集。数据集与代码已公开:https://HitEmotion.github.io/。
原文摘要 · Abstract (English)
Despite rapid progress in multimodal large language models (MLLMs), their capability for deep emotional understanding remains limited. We argue that genuine affective intelligence requires explicit modeling of Theory of Mind (ToM), the cognitive substrate from which emotions arise. To this end, we introduce HitEmotion, a ToM-grounded hierarchical benchmark that diagnoses capability breakpoints across increasing levels of cognitive depth. Second, we propose a ToM-guided reasoning chain that tracks mental states and calibrates cross-modal evidence to achieve faithful emotional reasoning. We further introduce TMPO, a reinforcement learning method that uses intermediate mental states as process-level supervision to guide and strengthen model reasoning. Extensive experiments show that HitEmotion exposes deep emotional reasoning deficits in state-of-the-art models, especially on cognitively demanding tasks. In evaluation, the ToM-guided reasoning chain and TMPO improve end-task accuracy and yield more faithful, more coherent rationales. In conclusion, our work provides the research community with a practical toolkit for evaluating and enhancing the cognition-based emotional understanding capabilities of MLLMs. Our dataset and code are available at: https://HitEmotion.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。