用视觉语言关联度优化生成,减少大模型幻觉
Grounding Language with Vision: A Conditional Mutual Information Calibrated Decoding Strategy for Reducing Hallucinations in LVLMs
- 基于条件互信息动态调整图文生成相关性
- 在多个基准上显著降低幻觉率,保持高效解码
- 适合需要高图文一致性任务的开发者使用
大型视觉语言模型(LVLMs)易产生幻觉,即生成内容看似合理却与输入图像无关。研究表明,这主要源于模型在解码时过度依赖语言先验而忽略视觉信息。为此,本文提出一种新的条件点互信息(C-PMI)校准解码策略,自适应增强生成文本与输入图像之间的相互依赖关系,以缓解幻觉。不同于仅关注文本词元采样的方法,我们联合建模视觉与文本词元对C-PMI的贡献,将幻觉抑制建模为双层优化问题,旨在最大化互信息。为求解该问题,设计了一种词元净化机制,通过采样与给定图像最相关的文本词元,同时优化生成响应中最相关的图像词元,动态调节解码过程。大量实验表明,该方法在多个基准上显著降低LVLM幻觉率,同时保持解码效率。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) are susceptible to hallucinations, where generated responses seem semantically plausible yet exhibit little or no relevance to the input image. Previous studies reveal that this issue primarily stems from LVLMs' over-reliance on language priors while disregarding the visual information during decoding. To alleviate this issue, we introduce a novel Conditional Pointwise Mutual Information (C-PMI) calibrated decoding strategy, which adaptively strengthens the mutual dependency between generated texts and input images to mitigate hallucinations. Unlike existing methods solely focusing on text token sampling, we propose to jointly model the contributions of visual and textual tokens to C-PMI, formulating hallucination mitigation as a bi-level optimization problem aimed at maximizing mutual information. To solve it, we design a token purification mechanism that dynamically regulates the decoding process by sampling text tokens remaining maximally relevant to the given image, while simultaneously refining image tokens most pertinent to the generated response. Extensive experiments across various benchmarks reveal that the proposed method significantly reduces hallucinations in LVLMs while preserving decoding efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。