arXiv:2410.12787cs.CV2024-10NeurIPS被引 55

首次系统评估多模态大模型在语言、视觉、音频中的幻觉问题

The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

  • 构建跨语言、视觉、音频三模态的幻觉评估基准CMM
  • 发现模型过度依赖单模态先验和虚假跨模态关联是主因
  • 适合关注多模态模型可靠性与幻觉缓解的研究者

近年来,大型多模态模型(LMMs)在多种任务上性能显著提升,研究持续探索视频、音频等新模态的融合。然而,现有LMMs仍易产生幻觉——即生成文本与真实多模态输入不一致。本文首次系统研究了涉及语言、视觉和音频三种主要模态的幻觉现象。研究揭示两大关键成因:对单模态先验的过度依赖及虚假的跨模态相关性。为此,我们提出基准测试《多模态诅咒》(The Curse of Multi-Modalities, CMM),全面评估LMMs中的幻觉问题,并深入分析其根本原因。研究发现,模态融合失衡与训练数据偏见是主要弱点,凸显了平衡跨模态学习与强化幻觉缓解策略的必要性。基于观察结果,我们提出若干潜在研究方向,以提升LMMs的可靠性。

原文摘要 · Abstract (English)

Recent advancements in large multimodal models (LMMs) have significantly enhanced performance across diverse tasks, with ongoing efforts to further integrate additional modalities such as video and audio. However, most existing LMMs remain vulnerable to hallucinations, the discrepancy between the factual multimodal input and the generated textual output, which has limited their applicability in various real-world scenarios. This paper presents the first systematic investigation of hallucinations in LMMs involving the three most common modalities: language, visual, and audio. Our study reveals two key contributors to hallucinations: overreliance on unimodal priors and spurious inter-modality correlations. To address these challenges, we introduce the benchmark The Curse of Multi-Modalities (CMM), which comprehensively evaluates hallucinations in LMMs, providing a detailed analysis of their underlying issues. Our findings highlight key vulnerabilities, including imbalances in modality integration and biases from training data, underscoring the need for balanced cross-modal learning and enhanced hallucination mitigation strategies. Based on our observations and findings, we suggest potential research directions that could enhance the reliability of LMMs.

多模态幻觉检测模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。