arXiv:2605.20772cs.CV2026-05中稿 · MICCAI 2026

通过视觉干预检测医疗图像问答中的幻觉,提升诊断可靠性。

VIHD: Visual Intervention-based Hallucination Detection for Medical Visual Question Answering

论文配图:VIHD: Visual Intervention-based Hallucination Detection for Medical Visual Question Answering
图 1 · 摘自论文原文
  • 用视觉标记遮蔽定位关键解码层,分析跨模态依赖关系。
  • 在三个医学VQA数据集上,比现有方法平均提升6.3%检测准确率。
  • 适合关注医疗AI安全与可解释性的研究人员和开发者。

尽管医学多模态大语言模型在辅助诊断方面展现出潜力,但仍频繁生成看似合理却缺乏视觉依据的幻觉回答,威胁临床决策。现有基于响应的检测方法依赖对原始或扰动输入的不确定性估计或逻辑验证,但此类外部扰动常为启发式且不考虑解码过程中生成标记与视觉标记间的内部跨模态依赖。为此,我们提出VIHD(视觉干预型幻觉检测)方法,通过有目标的视觉标记遮蔽来校准语义熵,以更有效地检测幻觉。VIHD通过视觉依赖探查(VDP)定位视觉主导的解码层,利用标记遮蔽执行视觉干预解码(VID),并量化校准后的语义熵(CSE)作为可靠的幻觉信号。在三个医学VQA基准上使用两种医学MLLM进行的大量实验表明,VIHD持续优于当前最佳方法,凸显细粒度视觉依赖对幻觉检测的重要性。代码将公开于 https://github.com/Jiayi-Chen-AU/VIHD。

原文摘要 · Abstract (English)

While medical Multimodal Large Language Models (MLLMs) have shown promise in assisting diagnosis, they still frequently generate hallucinated responses that appear linguistically plausible but lack visual evidence. Such hallucinations pose risks to clinical decision-making and necessitate effective detection. Existing introspective detection methods primarily perform uncertainty estimation or logical verification by analyzing model responses conditioned on original or perturbed inputs. However, such external perturbations are often heuristic and context-agnostic, which overlooks the internal cross-modal dependency between generated tokens and related visual tokens during decoding. To address this issue, we propose VIHD, a Visual Intervention-based Hallucination Detection method that leverages targeted visual token masking to calibrate semantic entropy for more effective hallucination detection. VIHD locates visually dominant decoder layers via Visual Dependency Probing (VDP), executes Visual Intervention Decoding (VID) via token masking to calibrate the semantic distribution, and quantifies the resulting Calibrated Semantic Entropy (CSE) as a reliable hallucination signal. Extensive experiments on three medical VQA benchmarks with two medical MLLMs demonstrate that VIHD consistently outperforms state-of-the-art methods, underscoring the importance of fine-grained visual dependency for hallucination detection. The code will be available at https://github.com/Jiayi-Chen-AU/VIHD

医学AI幻觉检测多模态视觉干预

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。