arXiv:2410.03176cs.CVcs.AI2024-10EMNLP被引 21

发现CLIP模型自身会生成虚假物体,提出对抗性数据增强方法缓解此问题。

Investigating and Mitigating Object Hallucinations in Pretrained Vision-Language (CLIP) Models

  • 通过构建含幻觉的负样本,研究CLIP模型的幻觉来源。
  • 所提方法使CLIP幻觉率显著降低,且可作更优视觉编码器。
  • 适用于需减少视觉错觉的多模态系统,如图文生成与检索。

大视觉语言模型(LVLMs)虽表现优异,但存在物体幻觉严重问题。本文深入探究了以CLIP为骨干的模型中幻觉成因,发现即使在孤立状态下,CLIP仍易产生物体幻觉,表明该问题并非仅由视觉与语言模态交互引起。为此,我们提出一种反事实数据增强方法,通过构造多种幻觉类型的负样本进行训练。实验表明,该方法能有效抑制CLIP的物体幻觉,且增强后的模型可作为高质量视觉编码器,显著缓解下游LVLM中的幻觉问题。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) have achieved impressive performance, yet research has pointed out a serious issue with object hallucinations within these models. However, there is no clear conclusion as to which part of the model these hallucinations originate from. In this paper, we present an in-depth investigation into the object hallucination problem specifically within the CLIP model, which serves as the backbone for many state-of-the-art vision-language systems. We unveil that even in isolation, the CLIP model is prone to object hallucinations, suggesting that the hallucination problem is not solely due to the interaction between vision and language modalities. To address this, we propose a counterfactual data augmentation method by creating negative samples with a variety of hallucination issues. We demonstrate that our method can effectively mitigate object hallucinations for CLIP model, and we show the the enhanced model can be employed as a visual encoder, effectively alleviating the object hallucination issue in LVLMs.

CLIP幻觉抑制多模态数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。