通过三层对比解码,让视觉语言模型回答更真实
Watermarking for Factuality: Guiding Vision-Language Models Toward Truth via Tri-layer Contrastive Decoding
- 用三层次解码结构区分成熟层、新手层和关键判断层
- 在多个基准测试中显著降低幻觉率,提升视觉准确性
- 无需训练,适合希望改进模型真实性的研究者
大型视觉语言模型(LVLM)在多模态任务上表现优异,甚至在某些情况下达到人类水平。然而,这些模型仍易产生幻觉——过度依赖单一模态或机械记忆训练数据而缺乏对输出的视觉真实锚定。为此,我们提出一种无需训练的三层对比解码水印方法,包含三个步骤:(1) 在解码层中选出一个成熟层与一个初学者层;(2) 通过一个与水印相关的提问识别出一个关键判断层,以评估该层是否具有良好的视觉依据;(3) 应用三层对比解码生成最终输出。在POPE、MME和AMBER等公开基准上的实验表明,该方法在减少LVLM幻觉方面达到当前最优性能,并生成更具视觉真实性的回答。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) have recently shown promising results on various multimodal tasks, even achieving human-comparable performance in certain cases. Nevertheless, LVLMs remain prone to hallucinations -- they often rely heavily on a single modality or memorize training data without properly grounding their outputs. To address this, we propose a training-free, tri-layer contrastive decoding with watermarking, which proceeds in three steps: (1) select a mature layer and an amateur layer among the decoding layers, (2) identify a pivot layer using a watermark-related question to assess whether the layer is visually well-grounded, and (3) apply tri-layer contrastive decoding to generate the final output. Experiments on public benchmarks such as POPE, MME and AMBER demonstrate that our method achieves state-of-the-art performance in reducing hallucinations in LVLMs and generates more visually grounded responses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。