无需标注,通过解剖区域反事实干预提升医学模型生成准确性
Counterfactual Anatomy-guided Spatial-Temporal Decoding for Annotation-Free Hallucination Mitigation in Medical VLMs

- 基于反事实干预自动识别与查询相关的解剖区域
- 在三个医学视觉语言模型上减少幻觉,优于依赖真值标注的方法
- 适合需要高可靠性医疗生成的临床应用与研究者
医学视觉语言模型(Med-VLMs)在医学视觉问答任务中表现优异,但仍易产生缺乏图像依据的幻觉性陈述。现有解码阶段的缓解方法通常缺乏解剖意识或依赖真实标注,限制了实用性。我们提出反事实解剖引导的空间-时间解码框架(CAST),完全在推理阶段运行,无需人工标注即可实现解剖层面的幻觉抑制。CAST通过广泛的医学分割自动发现与查询相关的解剖区域,并利用遮蔽后答案似然下降情况,通过反事实干预选择紧凑且因果信息丰富的区域。在此区域引导下,采用统一的对比解码过程,结合无分类器指导纠正空间注意力,以及逐步时间对比调节生成动态。在SLAKE和MIMIC-CXR数据集上对三种Med-VLMs的实验表明,CAST持续优于强基线,超越依赖真值标注的解码策略。结果表明,自动选取的紧凑区域可提供高效对比引导,无需专家标注,为提升空间定位准确性和降低幻觉提供了实用且通用的解决方案。代码已公开于https://github.com/csyifan/CAST。
原文摘要 · Abstract (English)
Medical vision-language models (Med-VLMs) have demonstrated strong performance on medical visual question answering, yet they remain prone to hallucination, generating clinically unsupported statements that are insufficiently grounded in image evidence. Mitigation methods applied during decoding offer a practical solution, but they typically lack anatomical awareness or rely heavily on ground truth annotations, which limits their applicability. We propose Counterfactual Anatomy-guided Spatial-Temporal decoding (CAST), a framework that operates entirely during inference and requires no manual annotations for anatomically grounded hallucination mitigation. CAST automatically discovers anatomical regions relevant to the given query through broad medical segmentation. It then selects a compact, causally informative area using counterfactual intervention based on the drop in answer likelihood under occlusion. Guided by this chosen region, CAST performs a unified contrastive decoding process, combining classifier-free guidance to correct spatial attention with stepwise temporal contrast to regulate generation dynamics. Experiments on the SLAKE and MIMIC-CXR datasets across three Med-VLMs demonstrate that CAST consistently outperforms strong baselines and surpasses decoding strategies reliant on ground truth. Our results indicate that compact, automatically selected regions provide highly effective contrastive guidance without expert annotations, offering a practical and generalizable solution for improving spatial grounding and reducing hallucinations. Code is available at https://github.com/csyifan/CAST.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。