用在线自校准减少视觉语言模型幻觉,提升生成真实性。
Online Self-Calibration Against Hallucination in Vision-Language Models

- 通过蒙特卡洛树搜索与双粒度奖励机制生成自监督偏好数据
- 在多个幻觉评测集上达到当前最优性能,且提升多模态通用能力
- 适合关注模型可靠性与生成真实性的研究人员
大型视觉语言模型常出现幻觉,即生成图像中不存在的视觉细节。现有偏好对齐方法依赖更强模型(如GPT)提供的离线监督,但存在「监督-感知错配」问题:学生模型被要求学习超出其感知能力的细粒度信息,导致猜解而非真实观察。为获取可靠的在线自监督信号,我们发现视觉语言模型在判别式验证任务中的准确率高于开放式生成任务,存在生成-判别差距。基于此,提出在线自校准框架OSCAR,结合蒙特卡洛树搜索与双粒度奖励机制构建偏好数据,并通过直接偏好优化迭代优化模型。大量实验表明,OSCAR在幻觉评测基准上表现领先,同时提升模型整体多模态能力。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) often suffer from hallucinations, generating descriptions that include visual details absent from the input image. Recent preference alignment methods typically rely on supervision distilled from stronger models such as GPT. However, this offline paradigm introduces a Supervision-Perception Mismatch: the student model is forced to align with fine-grained details beyond its perceptual capacity, learning to guess rather than to see. To obtain reliable self-supervision for online learning, we identify a Generative-Discriminative Gap within LVLMs, where models exhibit higher accuracy on discriminative verification than open-ended generation. Leveraging this capability, we propose \textbf{O}nline \textbf{S}elf-\textbf{CA}lib\textbf{R}ation (OSCAR), a framework that integrates Monte Carlo Tree Search with a Dual-Granularity Reward Mechanism to construct preference data and iteratively refines the model via Direct Preference Optimization. Extensive experiments demonstrate that OSCAR achieves state-of-the-art performance on hallucination benchmarks while improving general multimodal capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。