arXiv:2412.06775cs.CVcs.AI2024-12被引 11

通过视觉对比解码减少大模型幻觉,用不同图像处理方式提升准确性。

Delve into Visual Contrastive Decoding for Hallucination Mitigation of Large Vision-Language Models

  • 用图像降采样和编辑生成对比样本,增强模型对视觉内容的敏感性。
  • 不同对比样本在多个模型和数据集上效果差异显著,需针对性选择。
  • 提出融合多种对比样本的方法,提升跨场景适用性和鲁棒性。

大型视觉语言模型虽能生成与视觉输入相关的内容,但仍存在幻觉问题。现有方法采用对比解码,通过对比原始图像与视觉扰动样本的输出分布来缓解幻觉,且无需训练。本文探索了图像降采样与编辑等方法生成视觉对比样本,前者降低细节信息,后者引入新内容。进一步分析熵与分布距离等概率级指标发现,不同样本在不同模型与基准上的效果差异明显。基于此,提出一种简单有效的多样本融合策略,在多个基准上验证其有效性。

原文摘要 · Abstract (English)

While large vision-language models (LVLMs) have shown impressive capabilities in generating plausible responses correlated with input visual contents, they still suffer from hallucinations, where the generated text inaccurately reflects visual contents. To address this, recent approaches apply contrastive decoding to calibrate the model's response via contrasting output distributions with original and visually distorted samples, demonstrating promising hallucination mitigation in a training-free manner. However, the potential of changing information in visual inputs is not well-explored, so a deeper investigation into the behaviors of visual contrastive decoding is of great interest. In this paper, we first explore various methods for contrastive decoding to change visual contents, including image downsampling and editing. Downsampling images reduces the detailed textual information while editing yields new contents in images, providing new aspects as visual contrastive samples. To further study benefits by using different contrastive samples, we analyze probability-level metrics, including entropy and distribution distance. Interestingly, the effect of these samples in mitigating hallucinations varies a lot across LVLMs and benchmarks. Based on our analysis, we propose a simple yet effective method to combine contrastive samples, offering a practical solution for applying contrastive decoding across various scenarios. Extensive experiments are conducted to validate the proposed fusion method among different benchmarks.

视觉语言模型幻觉抑制对比解码图像编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。