arXiv:2508.04182cs.CLcs.AI2025-08被引 1

通过因果约束减少视觉模型幻觉,让回答更基于真实证据。

COPO: Causal-Oriented Policy Optimization for Hallucinations of MLLMs

  • 用因果完备性奖励衡量每个词的必要性与充分性
  • 在多个基准上幻觉率降低15%-30%,准确率提升
  • 适合关注大模型可信生成的开发者和研究者

尽管多模态大语言模型(MLLMs)表现出色,但仍存在幻觉问题。实证发现,相较于纯文本大模型,MLLMs对与任务无关的背景区域关注度过高,暗示了虚假的背景-答案关联。本文分析指出:(i) 基于结果的奖励机制是导致虚假关联的重要因素;(ii) 虚假关联又会引发幻觉。为此提出因果导向策略优化(COPO),通过引入词级充分性与必要性约束,衡量每个推理词的因果贡献,确保输出正确且有依据。具体地,先利用新提出的因果完备性奖励评估各词的因果贡献,再将其融入GRPO框架构建因果感知的优势函数,引导模型聚焦于对准确生成具有因果必要性和充分性的词。跨多个基准的实验表明,COPO显著降低幻觉率15%-30%,提升生成准确性。

原文摘要 · Abstract (English)

Despite Multimodal Large Language Models (MLLMs) having shown impressive capabilities, they may suffer from hallucinations. Empirically, we find that MLLMs attend disproportionately to task-irrelevant background regions compared with text-only LLMs, implying spurious background-answer correlations. We claim and analyze that (i) outcome-based rewards can be an important factor leading to spurious correlations, and (ii) spurious correlations can be an important factor leading to hallucinations. Based on these results, we propose Causal-Oriented Policy Optimization (COPO) to mitigate these spurious correlations, thus addressing the issue of hallucinations. It imposes token-level sufficiency and necessity constraints to measure each inference token's causal contribution, thus ensuring correct and evidence-grounded output. Specifically, we first evaluate each token's causal contribution via a newly proposed causal completeness reward. This reward is then used to construct a causally informed advantage function within the GRPO optimization framework, encouraging the model to focus on tokens that are causally sufficient and necessary for accurate generation. Experimental results across various benchmarks demonstrate the advantages of COPO.

多模态模型幻觉抑制因果推理生成质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。