arXiv:2511.17254cs.CVcs.AI2025-11NeurIPS被引 10

提出统一干预框架,解决视觉语言模型幻觉问题。

Intervene-All-Paths: Unified Mitigation of LVLM Hallucinations across Alignment Formats

  • 基于变压器因果结构,识别图像到文本、文本到文本等多路径幻觉成因。
  • 针对不同问答格式,定位并干预关键幻觉注意力头,显著降低幻觉率。
  • 适用于生成与判别式任务,适合关注模型可靠性研究者。

尽管大型视觉语言模型(LVLM)在多种任务中表现优异,仍易产生幻觉。本文提出一种与变压器因果架构对齐的综合干预框架,整合不同干预路径对幻觉的影响。研究发现,幻觉并非源于单一因果路径,而是图像到输入文本、图像到输出文本及文本到文本路径间的相互作用所致。首次发现LVLM会根据问答对齐格式依赖不同路径。基于此,提出简单有效的策略,分别针对判别式与生成式格式,定位并干预各路径中的关键幻觉注意力头。在多个基准测试上的实验表明,该方法能持续降低各类对齐类型下的幻觉水平。

原文摘要 · Abstract (English)

Despite their impressive performance across a wide range of tasks, Large Vision-Language Models (LVLMs) remain prone to hallucination. In this study, we propose a comprehensive intervention framework aligned with the transformer's causal architecture in LVLMs, integrating the effects of different intervention paths on hallucination. We find that hallucinations in LVLMs do not arise from a single causal path, but rather from the interplay among image-to-input-text, image-to-output-text, and text-to-text pathways. For the first time, we also find that LVLMs rely on different pathways depending on the question-answer alignment format. Building on these insights, we propose simple yet effective methods to identify and intervene on critical hallucination heads within each pathway, tailored to discriminative and generative formats. Experiments across multiple benchmarks demonstrate that our approach consistently reduces hallucinations across diverse alignment types.

视觉语言模型幻觉抑制注意力机制多路径干预

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。