arXiv:2505.20569cs.CVcs.AI2025-05ACL被引 4

用检索视觉对比解码,减少大模型幻觉物体生成

Retrieval Visual Contrastive Decoding to Mitigate Object Hallucinations in Large Vision-Language Models

  • 在逻辑层融合正负图像,抑制幻觉
  • 无需训练,直接优化解码过程
  • 适合需高可靠视觉生成的场景

尽管大型视觉语言模型取得了显著进展,物体幻觉(Object Hallucination, OH)仍是持续存在的挑战。基于此前无需额外模型训练即可缓解该问题的对比解码研究,本文提出一种名为RVCD(Retrieval Visual Contrastive Decoding)的新方法,以进一步抑制幻觉现象。RVCD在逻辑层同时利用正样本和负样本图像,明确参考由AI生成、代表单一概念的图像。该方法在不增加训练成本的前提下,显著优于现有的基于解码的抑制策略。

原文摘要 · Abstract (English)

Despite significant advancements in Large Vision-Language Models, Object Hallucination (OH) remains a persistent challenge. Building upon prior studies on contrastive decoding that address this issue without requiring additional model training, we introduce RVCD (Retrieval Visual Contrastive Decoding), an advanced method to suppress OH. RVCD leverages both negative and positive images at the logit level, explicitly referencing AI-generated images designed to represent a single concept. Our approach demonstrates substantial improvements over existing decoding-based methods.

视觉语言模型幻觉抑制对比解码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。