构建多模态情感智能评测基准,揭示大模型在真实情感交互中的能力短板。
EmoBench-M: Benchmarking Emotional Intelligence for Multimodal Large Language Models
- 基于心理学理论设计三层次评估框架,覆盖情绪识别到社会复杂情境分析。
- 27个主流多模态大模型测试显示最优仅达70.5分,远低于人类水平。
- 开源数据集与代码,助力情感智能研究与模型优化。
随着多模态大语言模型(MLLMs)融入机器人系统与AI应用,赋予其情感智能(EI)能力对有效感知、理解并回应人类情感至关重要。现有静态文本或图文基准忽视真实互动的多模态复杂性及情感表达的动态上下文依赖性,难以评估MLLMs的情感智能。为此,我们提出EmoBench-M,一个基于心理学理论的系统性基准,涵盖13个评估场景,分为三个层级:基础情绪识别(FER)、对话情感理解(CEU)和社会复杂情感分析(SCEA)。在27个前沿MLLM上进行评估,结合任务特定指标与大模型评价,结果显示性能与人类水平存在显著差距。最佳模型Gemini-3.0-Pro和GPT-5.2分别获得70.5和66.5分。专用模型如AffectGPT在部分场景表现突出,但整体缺乏全面情感智能。EmoBench-M提供了覆盖多样情感情境的综合评估框架,揭示当前模型的优势与局限。所有资源(数据集、代码)已公开于https://emo-gml.github.io/,支持后续研究推进。
原文摘要 · Abstract (English)
With the integration of multimodal large language models (MLLMs) into robotic systems and AI applications, embedding emotional intelligence (EI) capabilities is essential for enabling these models to perceive, interpret, and respond to human emotions effectively in real-world scenarios. Existing static, text-based, or text-image benchmarks overlook the multimodal complexities of real interactions and fail to capture the dynamic, context-dependent nature of emotional expressions, rendering them inadequate for evaluating MLLMs' EI capabilities. To address these limitations, we introduce EmoBench-M, a systematic benchmark grounded in established psychological theories, designed to evaluate MLLMs across 13 evaluation scenarios spanning three hierarchical dimensions: foundational emotion recognition (FER), conversational emotion understanding (CEU), and socially complex emotion analysis (SCEA). Evaluation was conducted on 27 state-of-the-art MLLMs, using both objective task-specific metrics and LLM-based evaluation, revealing a substantial performance gap relative to human-level competence. Even the best performing models, Gemini-3.0-Pro and GPT-5.2, achieve the highest scores on EmoBench-M, 70.5 and 66.5 points respectively. Specialized models such as AffectGPT exhibit uneven performance across EmoBench-M, demonstrating strengths in certain scenarios but generally lacking comprehensive emotional intelligence. By providing a comprehensive, multimodal evaluation framework, EmoBench-M captures both the strengths and weaknesses of current MLLMs across diverse emotional contexts. All benchmark resources, including datasets and code, are publicly available at https://emo-gml.github.io/, facilitating further research and advancement in MLLM emotional intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。