arXiv:2606.28401cs.CVcs.LG2026-06

用视觉线索生成更可信的偏好数据,减少视觉语言模型幻觉。

Vision-driven Preference Synthesis for Mitigating Hallucinations in VLMs

论文配图:Vision-driven Preference Synthesis for Mitigating Hallucinations in VLMs
图 1 · 摘自论文原文
  • 基于图像中重复出现的物体线索构建偏好对,避免依赖语言先验。
  • 在保持模型原有输出分布的前提下,提升对图像信息的利用度。
  • 显著降低幻觉率,适合追求视觉准确性的多模态模型研究者。

视觉语言模型(VLM)在视觉理解任务中表现优异,但仍存在生成与图像不符内容的幻觉问题。偏好对的构造方式直接影响模型的视觉忠实性。现有方法存在两大缺陷:(a) 干预式方法常导致策略分布显著偏离;(b) 采样式方法在构建过程中未能充分使用视觉信息。本文提出 ViPSy(Vision-driven Preference Synthesis),一种兼具策略一致性与视觉一致性的偏好数据生成框架。第一阶段,从语义对齐的图像变体中提取重复出现的物体级视觉线索,使偏好构建依赖于视觉而非语言先验。第二阶段,将该线索作为条件,引导模型自身滚动生成候选答案,确保候选结果既贴近策略分布,又具备视觉依据。实验表明,使用 ViPSy 构建的偏好对进行对齐后,新模型在幻觉缓解方面达到新基准:在 AMBER 和 Object HalBench 上分别降低幻觉率 35.7% 和 24.5%。同时,在 MMStar、MMVP、CV-Bench 等通用视觉对齐基准上表现更优,并在语义分割与 ImageNet 线性探测任务中取得提升,验证了其增强模型视觉能力的有效性。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have shown strong performance in visual understanding, yet they still suffer from hallucinations, generating content that is not grounded in the image. Preference alignment is a promising approach to improve visual faithfulness, but its success depends heavily on how preference pairs are constructed. Existing methods exhibit two key limitations; (a) intervention-based methods often introduce significant deviation from the policy distribution, and (b) sampling-based methods often underuse visual information during the construction. In this paper, we propose ViPSy (Vision-driven Preference Synthesis), a framework for constructing preference data that are both policy-aligned and visually grounded. Our framework consists of two stages; in the first stage, ViPSy derives a visual cue from recurring object-level content across semantically aligned image variants, so preference construction can rely on visual information rather than language priors. In the second stage, ViPSy conditions the policy's own rollouts on this cue, allowing candidates to be guided by visually grounded content while staying close to the policy's response distribution. The resulting candidates remain close to the policy's response distribution while better leveraging visual information from the image. Experiments show that the resulting VLM, preference-aligned with ViPSy-constructed preference pairs, achieves a new state-of-the-art in hallucination mitigation. Compared with the previous state-of-the-art method, it reduces hallucination rates on AMBER and Object HalBench by 35.7% and 24.5%, respectively. The resulting model further improves on general visual grounding benchmarks, e.g., MMStar, MMVP, and CV-Bench, while also yielding gains in semantic segmentation and ImageNet linear probing, underscoring the effectiveness of our framework in enhancing the model's visual capabilities.

视觉语言模型幻觉缓解偏好对构建多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。