arXiv:2602.07017cs.CVcs.AI2026-02

用语言模型引导的扰动方法,让医学图像分割更准更快更可解释。

XAI-CLIP: ROI-Guided Perturbation Framework for Explainable Medical Image Segmentation in Multimodal Vision-Language Models

  • 利用多模态语言模型定位关键解剖区域,指导扰动方向。
  • 运行速度提升60%,分割精度(Dice)提高44.6%,解释质量(IoU)提升96.7%。
  • 适合临床部署,生成清晰、符合解剖结构的可解释性热图。

医学图像分割在临床工作中至关重要,有助于精准诊断、治疗规划和疾病监测。尽管基于Transformer的模型在性能上优于传统卷积架构,但其可解释性差仍是阻碍临床信任与应用的主要障碍。现有可解释人工智能(XAI)技术,包括基于梯度的显著性方法和基于扰动的方法,通常计算成本高、需多次前向传播,且常产生噪声大或解剖无关的解释。为此,我们提出XAI-CLIP,一种基于感兴趣区域(ROI)引导的扰动框架,利用多模态视觉-语言模型嵌入定位临床上有意义的解剖区域,并指导解释过程。通过融合语言信息驱动的区域定位与医学图像分割,实施针对性的区域感知扰动,该方法生成更清晰、边界分明的显著性图,同时大幅降低计算开销。在FLARE22和CHAOS数据集上的实验表明,XAI-CLIP相较传统扰动方法,运行时间减少60%,Dice分数提升44.6%,遮挡解释的交并比(IoU)提升96.7%。定性结果进一步证实,生成的归因图更干净、解剖一致性更高,伪影更少,表明将多模态视觉-语言表示融入扰动型XAI框架,显著提升了可解释性与效率,推动透明且可临床部署的医学图像分割系统发展。

原文摘要 · Abstract (English)

Medical image segmentation is a critical component of clinical workflows, enabling accurate diagnosis, treatment planning, and disease monitoring. However, despite the superior performance of transformer-based models over convolutional architectures, their limited interpretability remains a major obstacle to clinical trust and deployment. Existing explainable artificial intelligence (XAI) techniques, including gradient-based saliency methods and perturbation-based approaches, are often computationally expensive, require numerous forward passes, and frequently produce noisy or anatomically irrelevant explanations. To address these limitations, we propose XAI-CLIP, an ROI-guided perturbation framework that leverages multimodal vision-language model embeddings to localize clinically meaningful anatomical regions and guide the explanation process. By integrating language-informed region localization with medical image segmentation and applying targeted, region-aware perturbations, the proposed method generates clearer, boundary-aware saliency maps while substantially reducing computational overhead. Experiments conducted on the FLARE22 and CHAOS datasets demonstrate that XAI-CLIP achieves up to a 60\% reduction in runtime, a 44.6\% improvement in dice score, and a 96.7\% increase in Intersection-over-Union for occlusion-based explanations compared to conventional perturbation methods. Qualitative results further confirm cleaner and more anatomically consistent attribution maps with fewer artifacts, highlighting that the incorporation of multimodal vision-language representations into perturbation-based XAI frameworks significantly enhances both interpretability and efficiency, thereby enabling transparent and clinically deployable medical image segmentation systems.

医学图像可解释性多模态分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。