通过不确定性引导探索,让多模态大模型更精准纠正视觉幻觉。
Uncertainty-Aware Exploratory Direct Preference Optimization for Multimodal Large Language Models

- 基于令牌级认知不确定性识别模型弱点,主动探索改进。
- 在优选样本中加强视觉薄弱项学习,避免对有效知识过度惩罚。
- 适合需要提升视觉对齐精度的多模态模型训练场景。
直接偏好优化(DPO)在缓解多模态大语言模型(MLLMs)幻觉问题上表现优异,通过偏好对进行学习。其核心挑战在于如何将序列级偏好转化为细粒度的视觉保真度监督。为保护易产生幻觉的视觉相关标记,现有方法通常依据模型自评估的视觉敏感性信号分配训练重点。然而,这种敏感性由仍在训练中的模型估计,引入了自我参照偏差:强化已掌握的视觉线索,忽视难以感知但关键的细节,从而限制了更深层次对齐。本文提出一种不确定性感知探索式直接偏好优化(UE-DPO)方法,使模型能够发现自身认知缺陷并主动探索自我修正,以令牌级认知不确定性为指导。具体而言,我们首先量化模型在给定图像中未能准确锚定标记预测所引发的不确定性;随后,基于不确定性感知的探索强度,在优选样本中加强对视觉不足标记的学习压力,并减轻在劣选样本中对有益知识的过度惩罚。此外,我们提供了该方法的理论依据,并通过大量实验验证了其有效性与鲁棒性。
原文摘要 · Abstract (English)
Direct Preference Optimization (DPO) has proven to be an effective solution for mitigating hallucination in Multimodal Large Language Models (MLLMs) by learning from preference pairs. One of its key challenges lies in how to transfer the sequence-level preference into fine-grained supervision on visual fidelity. To safeguard vision-related tokens that are prone to hallucination, existing methods typically allocate training emphasis according to the model's self-assessed visual sensitivity signals. However, such sensitivity, estimated by a model still under training, introduces self-referential bias: reinforcing already well-learned visual cues while neglecting hard-to-perceive but critical details, thereby limiting deeper alignment. In this work, we propose an Uncertainty-aware Exploratory Direct Preference Optimization (UE-DPO) method for MLLMs, which enables the model to uncover its cognitive deficiencies and actively explore for self-correction, guided by token-level epistemic uncertainty. Specifically, we first quantify the uncertainty from the model's failure to ground token predictions in the given image. Then, based on an uncertainty-aware exploration intensity, we encourage more learning pressure on visually deficient tokens in preferred samples, and alleviate the over-penalization of beneficial knowledge in dispreferred samples. Further, we provide a theoretical justification for our method, and extensive experiments demonstrate its effectiveness and robustness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。