通过解剖区域引导,精准抑制医学多模态模型的幻觉问题。
Anatomical Region-Guided Contrastive Decoding: A Plug-and-Play Strategy for Mitigating Hallucinations in Medical VLMs
- 用解剖掩码指导三阶段对比解码,实现区域聚焦。
- 在胸部X光、脑MRI等数据集上显著降低幻觉率。
- 无需训练,可直接插入现有模型,适合临床部署。
医学视觉-语言模型(MedVLMs)在临床应用中潜力巨大,但其可靠性受幻觉问题制约——模型常脱离视觉证据,依赖文本先验生成答案。现有缓解策略存在局限:基于训练的方法需昂贵专家标注,难以扩展;无训练干预如对比解码虽数据高效,但采用全局非针对性修正,在复杂临床场景中效果不可靠。为此,我们提出解剖区域引导的对比解码(ARCD),一种即插即用策略,通过解剖掩码引导三层次对比解码过程,在词元、注意力和逻辑层动态重加权,有效引导模型关注指定区域,强化解剖理解并抑制错误输出。在胸部X光、CT、脑MRI及眼超声等多种数据集上的大量实验表明,该方法显著提升区域理解能力,减少幻觉,增强整体诊断准确率。
原文摘要 · Abstract (English)
Medical Vision-Language Models (MedVLMs) show immense promise in clinical applicability. However, their reliability is hindered by hallucinations, where models often fail to derive answers from visual evidence, instead relying on learned textual priors. Existing mitigation strategies for MedVLMs have distinct limitations: training-based methods rely on costly expert annotations, limiting scalability, while training-free interventions like contrastive decoding, though data-efficient, apply a global, untargeted correction whose effects in complex real-world clinical settings can be unreliable. To address these challenges, we introduce Anatomical Region-Guided Contrastive Decoding (ARCD), a plug-and-play strategy that mitigates hallucinations by providing targeted, region-specific guidance. Our module leverages an anatomical mask to direct a three-tiered contrastive decoding process. By dynamically re-weighting at the token, attention, and logits levels, it verifiably steers the model's focus onto specified regions, reinforcing anatomical understanding and suppressing factually incorrect outputs. Extensive experiments across diverse datasets, including chest X-ray, CT, brain MRI, and ocular ultrasound, demonstrate our method's effectiveness in improving regional understanding, reducing hallucinations, and enhancing overall diagnostic accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。