首个评估多模态大模型情绪幻觉的基准,揭示模型在情绪理解上的系统性偏差。
EmotionHallucer: Evaluating Emotion Hallucinations in Multimodal Large Language Models
- 构建双维度评估框架:基于情绪心理学与真实多模态感知
- 38个模型测试显示闭源模型表现更优,推理能力提升检测效果
- 提出PEP-MEK框架,平均提升9.90%幻觉检测准确率
情绪理解是关键但具挑战性的任务。近年来多模态大语言模型(MLLMs)在此领域取得显著进展,但常出现幻觉,生成无关或荒谬内容。据我们所知,尚无专门针对情绪幻觉的评估工作。本文提出EmotionHallucer,首个用于检测和分析MLLM情绪幻觉的基准。不同于人类依赖生物与社会学习的情绪理解,MLLM仅依赖数据驱动学习,缺乏内在情感直觉。借助情绪心理学知识,从情绪心理学知识与真实多模态感知两方面评估幻觉。采用对抗性二元问答(QA)框架,设计精心构造的基本与幻觉配对样本,评估模型情绪幻觉倾向。在38个LLMs和MLLMs上评估发现:i) 多数当前模型存在严重情绪幻觉问题;ii) 闭源模型在检测情绪幻觉上优于开源模型,推理能力带来额外优势;iii) 模型在情绪心理学知识方面表现优于多模态情绪感知。作为副产品,我们提出PEP-MEK框架,在选定模型上平均提升9.90%的情绪幻觉检测性能。资源将发布于https://github.com/xxtars/EmotionHallucer。
原文摘要 · Abstract (English)
Emotion understanding is a critical yet challenging task. Recent advances in Multimodal Large Language Models (MLLMs) have significantly enhanced their capabilities in this area. However, MLLMs often suffer from hallucinations, generating irrelevant or nonsensical content. To the best of our knowledge, despite the importance of this issue, there has been no dedicated effort to evaluate emotion-related hallucinations in MLLMs. In this work, we introduce EmotionHallucer, the first benchmark for detecting and analyzing emotion hallucinations in MLLMs. Unlike humans, whose emotion understanding stems from the interplay of biology and social learning, MLLMs rely solely on data-driven learning and lack innate emotional instincts. Fortunately, emotion psychology provides a solid foundation of knowledge about human emotions. Building on this, we assess emotion hallucinations from two dimensions: emotion psychology knowledge and real-world multimodal perception. To support robust evaluation, we utilize an adversarial binary question-answer (QA) framework, which employs carefully crafted basic and hallucinated pairs to assess the emotion hallucination tendencies of MLLMs. By evaluating 38 LLMs and MLLMs on EmotionHallucer, we reveal that: i) most current models exhibit substantial issues with emotion hallucinations; ii) closed-source models outperform open-source ones in detecting emotion hallucinations, and reasoning capability provides additional advantages; iii) existing models perform better in emotion psychology knowledge than in multimodal emotion perception. As a byproduct, these findings inspire us to propose the PEP-MEK framework, which yields an average improvement of 9.90% in emotion hallucination detection across selected models. Resources will be available at https://github.com/xxtars/EmotionHallucer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。