arXiv:2506.08391cs.CV2025-06ICML被引 23

通过选择性与对比解码,减少视觉语言模型的感知幻觉。

SECOND: Mitigating Perceptual Hallucination in Vision-Language Models via Selective and Contrastive Decoding

  • 分步选择多尺度视觉信息,模拟人类视觉聚焦机制。
  • 在多个基准上显著降低幻觉率,提升图像理解准确性。
  • 适合关注视觉推理准确性的研究者与应用开发者。

尽管视觉语言模型(VLMs)取得了显著进展,现有模型仍受限于物体幻觉这一关键挑战,影响视觉理解的准确性。为此,我们提出SECOND:选择性与对比解码,一种新方法,使VLM能够以对象为中心的方式有效利用多尺度视觉信息,更贴近人类视觉感知。SECOND逐步选择并整合多尺度视觉信息,促进对图像的更精确解读。通过迭代对比这些视觉信息,SECOND显著减少了感知幻觉,并在广泛基准测试中表现优于现有方法。理论分析与实验表明,多尺度应用在VLM中潜力巨大,跨尺度的优先级与对比策略优于现有方法。

原文摘要 · Abstract (English)

Despite significant advancements in Vision-Language Models (VLMs), the performance of existing VLMs remains hindered by object hallucination, a critical challenge to achieving accurate visual understanding. To address this issue, we propose SECOND: Selective and Contrastive Decoding, a novel approach that enables VLMs to effectively leverage multi-scale visual information with an object-centric manner, closely aligning with human visual perception. SECOND progressively selects and integrates multi-scale visual information, facilitating a more precise interpretation of images. By contrasting these visual information iteratively, SECOND significantly reduces perceptual hallucinations and outperforms a wide range of benchmarks. Our theoretical analysis and experiments highlight the largely unexplored potential of multi-scale application in VLMs, showing that prioritizing and contrasting across scales outperforms existing methods.

视觉语言模型幻觉抑制多尺度融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。