通过强化模型对视觉证据的信任,减少多模态幻觉。
Clearer Sight, Fewer Lies: Oriented Pickup Preference Optimization for Multimodal Hallucination Mitigation

- 让模型学会判断视觉证据强弱,而非仅依赖语言先验。
- 在多个基准上显著降低幻觉率,提升生成准确性。
- 适合需要高可信视觉理解的多模态应用开发者。
多模态大语言模型容易产生幻觉,因其生成偏好未能充分与视觉证据对齐,导致过度依赖语言先验而非真实视觉信息。本文观察到:当利用模型自身注意力显式增强与查询相关的视觉证据时,生成结果更准确,表明许多错误并非源于感知缺失,而是模型对已有注意力证据的信任不足。为此,我们提出面向证据强度的定向抓取偏好优化(OPPO),一种关注视觉证据强度而非仅输出质量的对齐目标。OPPO 通过对比同一正确回答在强、锚定、弱证据视角下的表现,将原始视觉偏好转化为有序的证据对齐。进一步结合细粒度的片段级与词元级正则化以稳定训练。理论分析表明,有序证据边界可保证局部视觉敏感性的下界为正。在幻觉检测与通用基准上的大量实验表明,OPPO 持续优于基线方法。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) are prone to hallucination as their generation preferences are insufficiently calibrated to visual evidence, causing them to fall back on linguistic priors, rather than faithful grounding. In this work, we start from an empirical observation: when query-relevant visual evidence is explicitly strengthened using the model's own attention, generation becomes more accurate, suggesting that many failures do not arise solely from missing perception, but from an insufficient tendency to trust the evidence the model has already attended to. Motivated by this finding, we propose Oriented Pickup Preference Optimization (\texttt{OPPO}), an evidence-aware alignment objective that learns preferences over the strength of visual evidence, rather than only response quality. Concretely, \texttt{OPPO} contrasts the same faithful response under stronger, anchored, weaker-evidence views, turning naive visual preference into ordered visual-evidence alignment. We further combine this objective with fine-grained span-level and token-level regularization to stabilize the training. Besides, we provide a theoretical analysis showing that ordered evidence margins induce a positive lower bound on local visual sensitivity. Extensive evaluations across hallucination and general-purpose benchmarks demonstrate that \texttt{OPPO} consistently outperforms baseline methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。