arXiv:2602.21441cs.LGcs.AI2026-02

通过因果干预减少多模态模型幻觉,提升生成内容真实性。

Causal Decoding for Hallucination-Resistant Multimodal Large Language Models

  • 在生成过程中施加因果干预,抑制虚假对象关联。
  • 在多个基准上显著降低幻觉率,保持描述质量。
  • 适合需要高可信度输出的视觉问答与图像描述任务。

多模态大语言模型(MLLMs)在视觉-语言任务中能生成详细响应,但仍易出现物体幻觉(即引入图像中不存在的对象),影响实际应用可靠性。以往方法多依赖启发式惩罚、事后修正或通用解码调整,未能直接干预触发幻觉的机制,效果有限。本文提出一种因果解码框架,在生成阶段实施针对性因果干预,以削弱虚假依赖关系。该方法通过重构解码动态,有效减少虚假对象词元,同时保持描述质量。在图像字幕和问答任务的多个基准上,该框架显著降低物体幻觉率,实现当前最优的忠实度表现,且未损害整体输出质量。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) deliver detailed responses on vision-language tasks, yet remain susceptible to object hallucination (introducing objects not present in the image), undermining reliability in practice. Prior efforts often rely on heuristic penalties, post-hoc correction, or generic decoding tweaks, which do not directly intervene in the mechanisms that trigger object hallucination and thus yield limited gains. To address this challenge, we propose a causal decoding framework that applies targeted causal interventions during generation to curb spurious object mentions. By reshaping the decoding dynamics to attenuate spurious dependencies, our approach reduces false object tokens while maintaining descriptive quality. Across captioning and QA benchmarks, our framework substantially lowers object-hallucination rates and achieves state-of-the-art faithfulness without degrading overall output quality.

多模态幻觉抑制因果推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。