首个系统评估多模态大模型情感智能的基准,揭示其真实水平与短板。
MME-Emotion: A Holistic Evaluation Benchmark for Emotional Intelligence in Multimodal Large Language Models
- 构建覆盖6000+视频片段的多场景情感评测集,含8类任务和问答对。
- 顶尖模型情感识别率仅39.3%,推理链准确率56.0%,整体表现有限。
- 适用于评估模型情感理解与因果推理能力,适合研究者与开发者参考。
多模态大语言模型(MLLMs)在情感计算领域取得突破性进展,展现出涌现的情感智能。然而现有情感评测基准仍存在局限:难以评估模型在不同场景下的泛化能力,也缺乏对情绪成因推理能力的衡量。为此,我们提出MME-Emotion——一个系统性评估MLLM情感理解与推理能力的基准,具备可扩展性、多样场景与统一评测协议。该基准包含超过6,000个精心筛选的视频片段,对应任务导向的问答对,覆盖广泛场景,形成八项情感任务。同时引入融合多种度量的综合评估体系,并通过多智能体框架进行分析。对20个先进MLLM的严格评估显示:当前最佳模型在情感识别上仅达39.3%得分,链式推理(CoT)得分56.0%;通用模型(如Gemini-2.5-Pro)依赖通用多模态理解能力,而专用模型(如R1-Omni)可通过领域特定微调获得相近表现。本工作为提升MLLM情感智能提供基础支撑。
原文摘要 · Abstract (English)
Recent advances in multimodal large language models (MLLMs) have catalyzed transformative progress in affective computing, enabling models to exhibit emergent emotional intelligence. Despite substantial methodological progress, current emotional benchmarks remain limited, as it is still unknown: (a) the generalization abilities of MLLMs across distinct scenarios, and (b) their reasoning capabilities to identify the triggering factors behind emotional states. To bridge these gaps, we present \textbf{MME-Emotion}, a systematic benchmark that assesses both emotional understanding and reasoning capabilities of MLLMs, enjoying \textit{scalable capacity}, \textit{diverse settings}, and \textit{unified protocols}. As the largest emotional intelligence benchmark for MLLMs, MME-Emotion contains over 6,000 curated video clips with task-specific questioning-answering (QA) pairs, spanning broad scenarios to formulate eight emotional tasks. It further incorporates a holistic evaluation suite with hybrid metrics for emotion recognition and reasoning, analyzed through a multi-agent system framework. Through a rigorous evaluation of 20 advanced MLLMs, we uncover both their strengths and limitations, yielding several key insights: \ding{182} Current MLLMs exhibit unsatisfactory emotional intelligence, with the best-performing model achieving only $39.3\%$ recognition score and $56.0\%$ Chain-of-Thought (CoT) score on our benchmark. \ding{183} Generalist models (\emph{e.g.}, Gemini-2.5-Pro) derive emotional intelligence from generalized multimodal understanding capabilities, while specialist models (\emph{e.g.}, R1-Omni) can achieve comparable performance through domain-specific post-training adaptation. By introducing MME-Emotion, we hope that it can serve as a foundation for advancing MLLMs' emotional intelligence in the future.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。