通过分阶段监督减少多模态大模型幻觉,提升推理可信度。
Groc-PO: Grounded Context Preference Optimization for Truthful Multimodal LLMs

- 在物体、上下文、推理三阶段设计分步偏好优化,精准定位错误源头。
- 在GCPD数据集上,幻觉率降低23%,推理准确率提升18%。
- 适合关注多模态模型可靠性与可解释性的研究者和开发者。
尽管多模态大语言模型(MLLMs)进展迅速,仍存在视觉幻觉、内容虚构和不忠实推理等问题,严重影响其可信度与实用性。基于人类偏好的对齐方法如直接偏好优化(DPO)被广泛采用,但多模态推理错误常跨阶段传播,最终答案错误往往源于早期的定位阶段失误。标准DPO通常仅在最终答案层面进行偏好优化,导致早期定位阶段的监督间接且非阶段化,难以抑制因定位漂移和上下文不一致引发的误差传播。为此,我们提出面向多模态模型的有基底上下文偏好优化(Groc-PO)框架,并构建了有基底上下文偏好数据集(GCPD),围绕物体定位、上下文定位和有基底推理三个阶段组织多阶段偏好样本,以捕捉基底上下文的形成、整合与使用过程。通过在多个基底阶段引入更明确的偏好监督,Groc-PO强化了上下文依赖推理,减轻了跨阶段误差传播。大量实验表明,相较于标准DPO及其他强基线,Groc-PO在幻觉抑制、忠实推理和整体可靠性方面均有提升,验证了更显式基底监督在可信多模态推理中的价值。
原文摘要 · Abstract (English)
Despite the rapid progress of Multimodal Large Language Models (MLLMs), they still suffer from untruthfulness issues, such as visual hallucinations, content fabrication, and unfaithful reasoning, which substantially undermine their faithfulness and practical utility. Alignment methods based on human preference, such as Direct Preference Optimization (DPO), have been widely adopted to address these issues. However, multimodal reasoning errors often propagate across stages, and final-answer errors can often be traced to mistakes in early grounding stages, yet standard DPO typically applies preference optimization at the final-answer level. This credit-assignment challenge means that supervision for early grounding stages is indirect rather than stage-specific, making it difficult to suppress error propagation arising from grounding drift and context inconsistency. To address this, we propose Grounded Context Preference Optimization (Groc-PO), a grounded preference optimization framework for MLLMs. We further construct the Grounded Context Preference Dataset (GCPD), organizing multi-stage preference samples around three stages of Object Grounding, Contextual Grounding, and Grounded Reasoning, to capture the formation, integration, and utilization of grounded context. By introducing more explicit preference supervision over multiple grounded stages, Groc-PO strengthens context-dependent reasoning and mitigates cross-stage error propagation. Extensive experiments show that, compared with standard DPO and other strong baselines, Groc-PO achieves improved performance in hallucination mitigation, faithful reasoning, and overall reliability, supporting the value of more explicit grounded supervision for trustworthy multimodal reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。