arXiv:2509.21950cs.CV2025-09中稿 · ICLR被引 4

提出新评估框架,让大模型更准判断图像情绪。

Customizing Visual Emotion Evaluation for MLLMs: An Open-vocabulary, Multifaceted, and Scalable Approach

  • 设计情绪陈述判断任务,支持开放词汇与多维度评估
  • 自动化构建情绪语句,减少人工标注成本
  • 发现顶尖模型仍不及人类,尤其在主观感知上

近年来,多模态大语言模型(MLLMs)在各类任务中表现卓越,但其图像情绪感知能力仍存争议,零样本场景下结果不一。我们认为这源于现有评估方法的局限:忽略合理回答、情感分类体系单一、忽视上下文因素以及标注成本高。为此,我们提出一种情绪陈述判断任务,克服上述问题,并设计自动化流水线,以极低人力成本构建以情绪为中心的陈述。系统评估主流MLLMs后发现,它们在情绪解读和情境化判断上表现较好,但在理解感知主观性方面仍有不足。与人类相比,即使是最优的GPT4o也存在显著差距,凸显未来改进方向。本研究构建了基础评估框架并完成全面评测,旨在推动MLLM情感智能发展。

原文摘要 · Abstract (English)

Recently, Multimodal Large Language Models (MLLMs) have achieved exceptional performance across diverse tasks, continually surpassing previous expectations regarding their capabilities. Nevertheless, their proficiency in perceiving emotions from images remains debated, with studies yielding divergent results in zero-shot scenarios. We argue that this inconsistency stems partly from constraints in existing evaluation methods, including the oversight of plausible responses, limited emotional taxonomies, neglect of contextual factors, and labor-intensive annotations. To facilitate customized visual emotion evaluation for MLLMs, we propose an Emotion Statement Judgment task that overcomes these constraints. Complementing this task, we devise an automated pipeline that efficiently constructs emotion-centric statements with minimal human effort. Through systematically evaluating prevailing MLLMs, our study showcases their stronger performance in emotion interpretation and context-based emotion judgment, while revealing relative limitations in comprehending perception subjectivity. When compared to humans, even top-performing MLLMs like GPT4o demonstrate remarkable performance gaps, underscoring key areas for future improvement. By developing a fundamental evaluation framework and conducting a comprehensive MLLM assessment, we hope this work contributes to advancing emotional intelligence in MLLMs. Project page: https://github.com/wdqqdw/MVEI.

情绪识别多模态评估框架大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。