arXiv:2607.21061cs.CV2026-07被引 1

构建新评估框架与模型,提升大模型理解视觉情感的能力。

MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement

论文配图:MVEI & EmObserver: Empowering MLLM-Oriented Visual Emotional Intelligence via Emotion Statement Judgement
图 1 · 摘自论文原文
  • 提出情绪陈述判断法,让模型聚焦真假判定而非自由生成
  • 建立涵盖46.2万样本的大规模情感评测集MVEI
  • 推出专用于视觉情感分析的EmObserver模型,性能领先

情感图像内容分析(AICA)旨在识别和理解视觉内容引发的情绪,是迈向通用人工智能(AGI)的关键步骤。尽管多模态大模型(MLLMs)快速发展,但对其视觉情感智能的系统性评估仍严重缺失。我们发现,传统AICA范式与MLLM开放、指令驱动的特性存在结构性不匹配,导致四大局限:忽略合理回答、情绪分类体系有限、忽视上下文因素、标注成本高。为此,我们提出情绪陈述判断(ESJ),在保持输入表达力的同时,将输出限制为判别性判断。进一步构建了高效标注流程INSETS,生成了包含46.2万样本的INSETS-462k数据集,并建立了覆盖情感极性、情绪解读、场景上下文和感知主观性的严谨评测基准MVEI。在此基础上,我们开发了基于ESJ优化的面向情感的MLLM——EmObserver,通过多阶段训练策略实现性能突破。在多个AICA基准上的实验表明,EmObserver具有更高的准确率、泛化能力和推理可信度。综合来看,本研究确立了ESJ的有效性、MVEI的全面性以及EmObserver作为先进基线的价值。代码将开源于:https://github.com/wdqqdw/EmObserver。

原文摘要 · Abstract (English)

Affective Image Content Analysis (AICA) aims to recognize and understand emotions elicited by visual content, representing an indispensable step toward Artificial General Intelligence (AGI). However, despite the rapid progress of Multimodal Large Language Models (MLLMs), systematic evaluation of their visual emotional intelligence remains largely absent from recent model releases. We attribute this gap to a structural mismatch between conventional AICA paradigms and the open-ended, instruction-driven nature of MLLMs, where further analysis reveals four major limitations: omission of plausible responses, limited emotion taxonomies, neglect of contextual factors, and labor-intensive annotation. To overcome these barriers, we introduce Emotion Statement Judgement (ESJ), a statement-verification formulation that preserves the expressiveness of the input space while constraining outputs to discriminative judgements. We further develop INSETS, a labor-efficient pipeline that instantiates ESJ at scale by constructing INSETS-462k and supporting MVEI, a rigorously refined benchmark spanning sentiment polarity, emotion interpretation, scene context, and perception subjectivity. Beyond evaluation, we build EmObserver, an emotion-oriented MLLM optimized on ESJ through an elaborate multi-stage recipe. Extensive evaluation of broad-spectrum MLLMs on MVEI reveals fine-grained insights into current artificial visual emotional intelligence, while experiments on multiple AICA benchmarks demonstrate the accuracy, generalization, and reasoning faithfulness of EmObserver. Collectively, these results establish ESJ as a practical formulation, MVEI as a comprehensive benchmark, and EmObserver as an advanced baseline for advancing MLLM-oriented visual emotional intelligence. Code will be released at: https://github.com/wdqqdw/EmObserver.

视觉情感多模态评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。