arXiv:2505.15963cs.CVcs.CL2025-05被引 2

用模型自生成的幻觉内容动态训练,让视觉语言模型更准

OViP: Online Vision-Language Preference Learning for VLM Hallucination

  • 基于模型自身幻觉输出构建对比学习数据
  • 幻觉减少37%且保留多模态能力
  • 适合想提升大模型可靠性的人看

大型视觉语言模型(LVLMs)仍易产生与视觉输入不符的幻觉。现有训练方法依赖预设或随机编辑的负样本,无法反映真实错误,限制了训练效果。本文提出在线视觉-语言偏好学习框架OViP,根据模型自身幻觉输出动态构建对比训练数据。通过识别响应对间的语义差异,并利用扩散模型合成负样本图像,OViP实现实时相关性更强的监督信号。该故障驱动的训练使文本与视觉偏好自适应对齐。此外,我们优化了评估协议,更好捕捉幻觉抑制与表达力之间的权衡。在幻觉和通用基准测试中,OViP不仅显著降低幻觉率,同时保持核心多模态能力,并大幅提升训练效率。代码已公开于https://github.com/lsjlsj35/Online-Vision-Language-Preference-Learning-for-VLM-Hallucination。

原文摘要 · Abstract (English)

Large vision-language models (LVLMs) remain vulnerable to hallucination, often generating content misaligned with visual inputs. Although recent training-based approaches aim to mitigate hallucination, they typically rely on predefined or randomly edited negative samples that do not reflect actual model errors, thus limiting training efficacy. In this work, we propose an Online Vision-language Preference Learning (OViP) framework that dynamically constructs contrastive training data based on the model's own hallucinated outputs. By identifying semantic differences between sampled response pairs and synthesizing negative images using a diffusion model, OViP generates more relevant supervision signals in real time. This failure-driven training enables adaptive alignment of both textual and visual preferences. Moreover, we refine existing evaluation protocols to better capture the trade-off between hallucination suppression and expressiveness. Experiments on hallucination and general benchmarks demonstrate that OViP not only reduces hallucinations while preserving core multi-modal capabilities, but also substantially improves training efficiency. Code is available at https://github.com/lsjlsj35/Online-Vision-Language-Preference-Learning-for-VLM-Hallucination.

视觉语言模型幻觉抑制在线学习扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。