arXiv:2606.03376cs.CVcs.AI2026-06被引 1

让视觉语言模型更准,减少幻觉,靠自生成偏好对训练

P$^2$-DPO: Grounding Hallucination in Perceptual Processing via Calibration Direct Preference Optimization

论文配图:P$^2$-DPO: Grounding Hallucination in Perceptual Processing via Calibration Direct Preference Optimization
图 1 · 摘自论文原文
  • 模型自动生成偏好对,聚焦注意力区域的感知优化
  • 在图像退化场景下仍保持高准确率,显著提升视觉鲁棒性
  • 无需人工标注,适合追求高效、真实场景表现的开发者

视觉语言大模型中的幻觉问题近年受到广泛关注。直接偏好优化(DPO)通过人类修正的偏好对进行学习,缓解了幻觉问题。然而,现有方法未针对性解决注意力区域的感知瓶颈,也未增强对图像退化的视觉鲁棒性。此外,偏好对常为视觉无关且具有固有的离策略特性,限制了其指导模型学习的效果。为此,本文提出感知处理直接偏好优化(P²-DPO),一种新训练范式:模型自生成偏好对,直接针对感知瓶颈并避免视觉无关与离策略数据问题。该方法包含:(1) 针对焦点增强感知与视觉鲁棒性的在线偏好对构建方法;(2) 精确校准视觉信号与文本因果生成的校准损失。实验表明,在相似训练数据量与成本下,P²-DPO 在多个基准上优于依赖昂贵人工反馈的强基线。注意力区域保真度(ARF)与图像退化场景评估验证了其在缓解注意力区域感知瓶颈与提升视觉鲁棒性方面的有效性。

原文摘要 · Abstract (English)

Hallucination has recently garnered significant research attention in Large Vision-Language Models (LVLMs). Direct Preference Optimization (DPO) aims to learn directly from the corrected preferences provided by humans, thereby addressing the hallucination issue. Despite its success, this paradigm has yet to specifically target the perceptual bottleneck in attended regions or address insufficient Visual Robustness against image degradation. Furthermore, existing preference pairs are often vision-agnostic and their inherently off-policy nature limits their effectiveness in guiding model learning. To address these challenges, we propose Perceptual Processing Direct Preference Optimization (P$^2$-DPO), a novel training paradigm in which the model generates and learns from its own preference pairs, thereby directly addressing the identified visual bottlenecks while inherently avoiding the issues of vision-agnostic and off-policy data. It introduces: (1) an on-policy preference pairs construction method targeting Focus-and-Enhance perception and Visual Robustness, and (2) a well-designed Calibration Loss to precisely align visual signals with the causal generation of text. Experimental results demonstrate that with a comparable amount of training data and cost, P$^2$-DPO outperforms strong baselines that rely on costly human feedback on benchmarks. Furthermore, evaluations on Attention Region Fidelity (ARF) and image degradation scenarios validate the effectiveness of P$^2$-DPO in addressing perceptual bottleneck in attended regions and improving Visual Robustness against degraded inputs.

视觉语言模型幻觉抑制偏好优化感知对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。