通过构建硬样本偏好对,减少多模态大模型的幻觉问题。
Grounded Post-Training with Hard Examples for Reducing Hallucination in Multimodal Large Language Models

- 基于失败边界构造幻觉相关偏好数据,针对性优化
- 在多个基准上幻觉率下降,视觉一致性显著提升
- 适合关注多模态生成可信度的研究者和开发者
幻觉仍是视觉语言模型(VLMs)的核心挑战,自回归生成可能因联合概率建模下的似然最大化,产生语法合理但物理不一致或视觉无依据的回答。本文提出一种分阶段偏好优化框架,通过有针对性的多模态数据构建来减少幻觉。不同于直接在通用指令跟随数据上优化,本方法逐步在已知失败边界附近构建聚焦幻觉的偏好对,强调模糊空间方位、物体关系、OCR不确定性及对抗性错误前提训练。通过最小扰动却视觉不一致的幻觉负例生成,使直接偏好优化(DPO)更有效区分有依据推理与可笑幻觉。在开源基准和真实多模态评估场景中的实验表明,该方法提升了视觉一致性,降低了幻觉率,并生成更丰富的有依据回答。跨模型定性评估进一步显示,该多模态LLM DPO框架在模糊空间推理和对抗性错误前提场景中,表现优于多个前沿专有VLMs。结果表明,幻觉不仅源于模型容量限制,也源于自回归概率生成在弱视觉支撑下倾向于选择语法合理延续的内在倾向。未来工作可探索物理一致性建模、不确定性感知的多模态推理及非标准自回归解码架构。
原文摘要 · Abstract (English)
Hallucination remains a fundamental challenge in vision-language models (VLMs), where autoregressive generation may produce linguistically plausible yet physically inconsistent or visually ungrounded responses due to likelihood maximization under joint probabilistic modeling. We propose a stage-wise preference optimization framework for hallucination reduction through targeted multimodal data construction. Rather than directly optimizing on generic instruction-following data, our approach progressively constructs hallucination-focused preference pairs near known failure boundaries. The framework emphasizes ambiguous spatial orientation, object relationships, OCR uncertainty, and adversarial false-premise training. Hallucinated negatives are generated through minimally perturbed yet visually inconsistent alternatives, enabling Direct Preference Optimization (DPO) to better separate grounded reasoning from plausible hallucination. Experiments on open-source benchmarks and real-world multimodal evaluation scenarios demonstrate improved grounding consistency, reduced hallucination, and more informative grounded responses. Cross-model qualitative evaluation further shows that the proposed multimodal LLM DPO framework produces more visually grounded responses than several frontier proprietary VLMs, such as in ambiguous spatial reasoning and adversarial false-premise settings. The results suggest that hallucination may arise not only from limited model capacity, but also from inherent tendencies of autoregressive probabilistic generation to favor linguistically plausible continuations under weak visual grounding. Future work may explore physical consistency modeling, uncertainty-aware multimodal reasoning, and architectural alternatives beyond standard autoregressive decoding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。