arXiv:2508.16859cs.CV2025-08被引 25

构建多轮多模态情绪理解与推理基准,推动机器从识别情绪迈向深度推理。

Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark

  • 设计多轮多模态情绪推理数据集,含1451段真实场景视频与5101个递进问题。
  • 现有大模型在复杂情绪推理任务上表现不佳,平均准确率低于40%。
  • 提出多智能体框架,分工处理背景、人物动态和事件细节,提升推理能力。

多模态大语言模型(MLLMs)凭借强大的感知与推理能力,在多个领域广泛应用。在心理学领域,这些模型有望实现对人类情绪与行为的深层理解。然而,当前研究主要集中于提升情绪识别能力,忽视了对情绪推理这一关键环节的探索,而情绪推理对提升人机交互的自然性与有效性至关重要。为此,本文提出一个包含1,451段真实生活场景视频、5,101个递进式问题的多轮多模态情绪理解与推理(MTMEUR)基准。问题涵盖情绪识别、情绪成因分析、未来行为预测等多个维度。同时,我们提出一种多智能体框架,各智能体分别专注于背景上下文、角色动态和事件细节等不同方面,以增强系统的综合推理能力。我们在该基准上对现有MLLMs及所提方法进行实验,结果表明,大多数模型在此任务上面临严峻挑战。

原文摘要 · Abstract (English)

Multimodal large language models (MLLMs) have been widely applied across various fields due to their powerful perceptual and reasoning capabilities. In the realm of psychology, these models hold promise for a deeper understanding of human emotions and behaviors. However, recent research primarily focuses on enhancing their emotion recognition abilities, leaving the substantial potential in emotion reasoning, which is crucial for improving the naturalness and effectiveness of human-machine interactions. Therefore, in this paper, we introduce a multi-turn multimodal emotion understanding and reasoning (MTMEUR) benchmark, which encompasses 1,451 video data from real-life scenarios, along with 5,101 progressive questions. These questions cover various aspects, including emotion recognition, potential causes of emotions, future action prediction, etc. Besides, we propose a multi-agent framework, where each agent specializes in a specific aspect, such as background context, character dynamics, and event details, to improve the system's reasoning capabilities. Furthermore, we conduct experiments with existing MLLMs and our agent-based method on the proposed benchmark, revealing that most models face significant challenges with this task.

多模态情绪推理大模型评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。