arXiv:2510.10466cs.CV2025-10被引 2

通过跨模态引导减少视觉语言模型的幻觉问题

When Images Speak Louder: Mitigating Language Bias-induced Hallucinations in VLMs through Cross-Modal Guidance

  • 不需训练,通过削弱视觉-语言注意力来抑制语言偏见
  • 在多个基准上显著降低幻觉率,提升生成内容相关性
  • 适合希望提升视觉理解准确性的研究者和开发者

视觉语言模型(VLMs)在多模态理解方面表现出色,但常出现幻觉现象:生成的文本虽语言流畅,却与图像内容无关。本文分析了语言偏见如何导致幻觉,并提出一种无需训练的解码方法——跨模态引导(CMG)。该方法通过自适应地遮蔽选定变压器层中最具影响力的图像标记注意力权重,人为破坏视觉-语言关联,从而增强对视觉上下文的感知,有效抑制语言偏见。实验表明,CMG在无额外条件或训练成本的情况下,在多个幻觉专项基准上均表现优异,且能泛化至不同VLM架构。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have shown solid ability for multimodal understanding of both visual and language contexts. However, existing VLMs often face severe challenges of hallucinations, meaning that VLMs tend to generate responses that are only fluent in the language but irrelevant to images in previous contexts. To address this issue, we analyze how language bias contributes to hallucinations and then introduce Cross-Modal Guidance(CMG), a training-free decoding method that addresses the hallucinations by leveraging the difference between the output distributions of the original model and the one with degraded visual-language attention. In practice, we adaptively mask the attention weight of the most influential image tokens in selected transformer layers to corrupt the visual-language perception as a concrete type of degradation. Such a degradation-induced decoding emphasizes the perception of visual contexts and therefore significantly reduces language bias without harming the ability of VLMs. In experiment sections, we conduct comprehensive studies. All results demonstrate the superior advantages of CMG with neither additional conditions nor training costs. We also quantitatively show CMG can improve different VLM's performance on hallucination-specific benchmarks and generalize effectively.

视觉语言模型幻觉抑制跨模态引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。