测试大模型看图理解情绪的能力,发现视觉识别不等于情感理解。
Seeing is Not Understanding: A Benchmark on Perception-Cognition Disparities in Large Language Models
- 构建分层任务框架,从看图识物到理解情绪意图
- 在350个真实社交图文样本上,模型情感理解准确率仅41%
- 适合关注多模态情感分析与模型认知局限的研究者
随着多模态大语言模型(MLLMs)的快速发展,其在多种视觉-语言任务中表现出色。然而,现有评估基准主要聚焦于客观视觉问答或描述生成,难以衡量模型对复杂、主观人类情绪的理解能力。为此,我们提出EmoBench-Reddit,一个全新的分层式多模态情绪理解基准。该数据集包含350个来自Reddit的精心筛选样本,每条数据包含一张图像、用户提供的文本及由用户标签确认的情绪类别(悲伤、幽默、讽刺、快乐)。我们设计了从基础感知到高级认知的分层任务框架,每个样本包含六道选择题和一道开放题,难度逐步提升。感知任务评估模型识别基本视觉元素(如颜色、物体)的能力,认知任务则要求进行场景推理、意图理解及结合文本语境的深层共情。通过Claude 4辅助与人工校验确保标注质量。我们对九个主流MLLMs(包括GPT-5、Gemini-2.5-pro、GPT-4o)进行了全面评估。
原文摘要 · Abstract (English)
With the rapid advancement of Multimodal Large Language Models (MLLMs), they have demonstrated exceptional capabilities across a variety of vision-language tasks. However, current evaluation benchmarks predominantly focus on objective visual question answering or captioning, inadequately assessing the models' ability to understand complex and subjective human emotions. To bridge this gap, we introduce EmoBench-Reddit, a novel, hierarchical benchmark for multimodal emotion understanding. The dataset comprises 350 meticulously curated samples from the social media platform Reddit, each containing an image, associated user-provided text, and an emotion category (sad, humor, sarcasm, happy) confirmed by user flairs. We designed a hierarchical task framework that progresses from basic perception to advanced cognition, with each data point featuring six multiple-choice questions and one open-ended question of increasing difficulty. Perception tasks evaluate the model's ability to identify basic visual elements (e.g., colors, objects), while cognition tasks require scene reasoning, intent understanding, and deep empathy integrating textual context. We ensured annotation quality through a combination of AI assistance (Claude 4) and manual verification.We conducted a comprehensive evaluation of nine leading MLLMs, including GPT-5, Gemini-2.5-pro, and GPT-4o, on EmoBench-Reddit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。