用无关图像对比,精准抑制语言先验,减少幻觉但不丢质量
Cross-Image Contrastive Decoding: Precise, Lossless Suppression of Language Priors in Large Vision-Language Models
- 用不同图像做对比,避免原图扰动带来的信号失真
- 动态选择抑制时机,降低幻觉率同时保持生成质量
- 适合需要高视觉一致性的图文生成任务
大型视觉语言模型(LVLMs)过度依赖语言先验是导致幻觉的主要原因,常产生语言通顺但视觉不符的输出。现有对比解码方法通过扰动原图构造对比视觉输入,导致对比分布失真、信号不完整,且过度抑制语言先验。我们观察到语言先验在不同图像间具有稳定性,提出跨图像对比解码(CICD),使用无关图像作为对比输入。为避免过度抑制语言先验影响生成质量,引入基于模型行为差异的动态选择机制,有选择地抑制语言先验。在多个基准和模型上的实验验证了CICD的有效性与泛化性,尤其在图像描述任务中表现突出。
原文摘要 · Abstract (English)
Over-reliance on language priors is a major cause of hallucinations in Large Vision-Language Models (LVLMs), often leading to outputs that are linguistically plausible but visually inconsistent. Recent studies have explored contrastive decoding as a training-free solution. However, these methods typically construct contrastive visual inputs by perturbing the original image, resulting in distorted contrastive distributions, incomplete contrastive signals, and excessive suppression of language priors. Motivated by the observation that language priors tend to remain consistent across different images, we propose Cross-Image Contrastive Decoding (CICD), a simple yet effective training-free method that uses unrelated images as contrastive visual inputs. To address the issue of over-suppressing language priors, which can negatively affect the quality of generated responses, we further introduce a dynamic selection mechanism based on the cross-image differences in model behavior. By selectively suppressing language priors, our method reduces hallucinations without compromising the model's performance. Extensive experiments across multiple benchmarks and LVLMs confirm the effectiveness and generalizability of CICD, particularly in image captioning, where language priors are especially dominant.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。