arXiv:2601.05159cs.CVcs.AI2026-01ACL被引 11

通过可解释的双向因果调控,降低多模态大模型的幻觉概率。

Vision-Language Introspection: Mitigating Overconfident Hallucinations in MLLMs via Interpretable Bi-Causal Steering

  • 基于概率冲突检测诊断幻觉风险,定位视觉关键证据。
  • 动态分离视觉信息与背景噪声,提升推理准确性。
  • 无需训练,适用于主流多模态大模型,效果显著。

物体幻觉严重威胁多模态大语言模型的可靠性,根源在于认知内省失败,模型盲目信任语言先验而非具体视觉证据。现有缓解方法受限:对比解码仅表面处理,未修复内部语义错位;现有潜在空间调制依赖静态向量,缺乏实例级精度。本文提出无需训练的推理框架Vision-Language Introspection(VLI),模拟元认知自我修正过程。VLI首先通过属性内省进行幻觉风险诊断,利用概率冲突检测定位因果视觉锚点;随后采用可解释双向因果调制,动态隔离视觉证据并抑制盲信,实现自适应校准。在先进模型上取得领先性能,在MMHal-Bench上将物体幻觉率降低12.67%,在POPE上准确率提升5.8%。

原文摘要 · Abstract (English)

Object hallucination critically undermines the reliability of Multimodal Large Language Models, often stemming from a fundamental failure in cognitive introspection, where models blindly trust linguistic priors over specific visual evidence. Existing mitigations remain limited: contrastive decoding approaches operate superficially without rectifying internal semantic misalignments, while current latent steering methods rely on static vectors that lack instance-specific precision. We introduce Vision-Language Introspection (VLI), a training-free inference framework that simulates a metacognitive self-correction process. VLI first performs Attributive Introspection to diagnose hallucination risks via probabilistic conflict detection and localize the causal visual anchors. It then employs Interpretable Bi-Causal Steering to actively modulate the inference process, dynamically isolating visual evidence from background noise while neutralizing blind confidence through adaptive calibration. VLI achieves state-of-the-art performance on advanced models, reducing object hallucination rates by 12.67% on MMHal-Bench and improving accuracy by 5.8% on POPE.

多模态幻觉抑制可解释性推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。