用自校正机制让视觉语言模型自己纠错,减少幻觉。
Aligning with Your Own Voice: Self-Corrected Preference Learning for Hallucination Mitigation in LVLMs

- 模型通过自我诊断和验证生成符合自身分布的训练数据
- 仅用5200条数据就显著降低幻觉率
- 适合需要低成本高精度去幻觉的场景
大型视觉语言模型(LVLMs)常出现幻觉问题。现有基于偏好学习的方法多依赖专有模型构建偏好数据集,导致专有模型与目标模型间分布不匹配,影响对齐效率。为此,我们提出一种基于自校正验证的偏好学习框架(AVES-DPO),利用模型内在知识生成分布一致的数据。该方法采用基于共识的验证机制识别多样幻觉,并引导模型自我修正,从而生成严格符合其内部分布的偏好对。大量实验表明,AVES-DPO在减少幻觉方面优于现有基线,且仅需5200个样本即可实现有效对齐。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) frequently suffer from hallucinations. Existing preference learning-based approaches largely rely on proprietary models to construct preference datasets. We identify that this reliance introduces a distributional mismatch between the proprietary and target models that hinders efficient alignment. To address this, we propose Alignment via VErified Self-correction DPO (AVES-DPO), a framework that aligns LVLMs using in-distribution data derived from the model's intrinsic knowledge. Our approach employs a consensus-based verification mechanism to diagnose diverse hallucinations and guides the model to self-correct, thereby generating preference pairs strictly compatible with its internal distribution. Extensive experiments demonstrate that AVES-DPO surpasses existing baselines in hallucination mitigation while requiring only 5.2k samples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。