arXiv:2504.07521cs.AIcs.MM2025-04CVPR被引 10

让AI理解情绪背后的因果,不止识别情绪本身

Why We Feel: Breaking Boundaries in Emotional Reasoning with Multimodal Large Language Models

  • 提出情绪解释任务(EI),聚焦情绪触发的显性与隐性原因
  • 构建含1615个基础样本和50个复杂样本的EIBench基准
  • 设计自问自答式标注流程,提升多模态模型推理质量

现有情绪分析多关注情绪类型(如快乐、悲伤、愤怒),却忽视其深层成因。本文提出情绪解释(Emotion Interpretation, EI),聚焦驱动情绪反应的因果因素——包括显性(如可见物体、人际互动)和隐性(如文化背景、未呈现事件)。与传统情绪识别不同,EI需推理情绪触发源而非简单分类。为此,我们构建了大规模基准EIBench,包含1,615个基础样本和50个复杂样本,每例均需基于理由的解释。提出粗到细自问自答(CFSA)标注流程,引导视觉语言模型通过迭代问答生成高质量标签。在开源与专有大模型上四组实验表明,复杂场景下性能差距显著,凸显EI在增强共情与上下文感知型AI应用中的潜力。相关数据集与方法已公开于https://github.com/Lum1104/EIBench,为多模态因果分析与下一代情感计算提供基础。

原文摘要 · Abstract (English)

Most existing emotion analysis emphasizes which emotion arises (e.g., happy, sad, angry) but neglects the deeper why. We propose Emotion Interpretation (EI), focusing on causal factors-whether explicit (e.g., observable objects, interpersonal interactions) or implicit (e.g., cultural context, off-screen events)-that drive emotional responses. Unlike traditional emotion recognition, EI tasks require reasoning about triggers instead of mere labeling. To facilitate EI research, we present EIBench, a large-scale benchmark encompassing 1,615 basic EI samples and 50 complex EI samples featuring multifaceted emotions. Each instance demands rationale-based explanations rather than straightforward categorization. We further propose a Coarse-to-Fine Self-Ask (CFSA) annotation pipeline, which guides Vision-Language Models (VLLMs) through iterative question-answer rounds to yield high-quality labels at scale. Extensive evaluations on open-source and proprietary large language models under four experimental settings reveal consistent performance gaps-especially for more intricate scenarios-underscoring EI's potential to enrich empathetic, context-aware AI applications. Our benchmark and methods are publicly available at: https://github.com/Lum1104/EIBench, offering a foundation for advanced multimodal causal analysis and next-generation affective computing.

情绪分析多模态因果推理大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。