提出可稳定解释的医学影像诊断模型,兼顾准确与可信度。
Stable Vision Concept Transformers for Medical Diagnosis
- 融合图像特征与概念特征,提升决策能力
- 在四组医疗数据集上保持高精度且可解释
- 抗干扰强,适合临床对可靠解释的需求
可解释性是医疗领域的重要关切,促使研究者探索可解释AI(XAI)。其中,概念瓶颈模型(CBMs)通过引入人类可理解的高层概念层来限制模型隐空间,受到广泛关注。然而现有方法仅依赖概念特征做预测,忽视了医学图像中的原始特征嵌入,导致与原模型性能存在差距。为此,本文提出视觉概念变压器(VCT),并在此基础上进一步提出稳定视觉概念变压器(SVCT),以视觉变换器(ViT)为骨干网络,结合概念层。SVCT通过融合概念特征与图像特征增强决策能力,并利用去噪扩散平滑(Denoised Diffusion Smoothing)确保模型输出的忠实性。在四个医疗数据集上的全面实验表明,相比基线模型,本方法在保持准确率的同时具备可解释性;即使面对输入扰动,SVCT仍能持续提供可信解释,满足医疗应用需求。
原文摘要 · Abstract (English)
Transparency is a paramount concern in the medical field, prompting researchers to delve into the realm of explainable AI (XAI). Among these XAI methods, Concept Bottleneck Models (CBMs) aim to restrict the model's latent space to human-understandable high-level concepts by generating a conceptual layer for extracting conceptual features, which has drawn much attention recently. However, existing methods rely solely on concept features to determine the model's predictions, which overlook the intrinsic feature embeddings within medical images. To address this utility gap between the original models and concept-based models, we propose Vision Concept Transformer (VCT). Furthermore, despite their benefits, CBMs have been found to negatively impact model performance and fail to provide stable explanations when faced with input perturbations, which limits their application in the medical field. To address this faithfulness issue, this paper further proposes the Stable Vision Concept Transformer (SVCT) based on VCT, which leverages the vision transformer (ViT) as its backbone and incorporates a conceptual layer. SVCT employs conceptual features to enhance decision-making capabilities by fusing them with image features and ensures model faithfulness through the integration of Denoised Diffusion Smoothing. Comprehensive experiments on four medical datasets demonstrate that our VCT and SVCT maintain accuracy while remaining interpretable compared to baselines. Furthermore, even when subjected to perturbations, our SVCT model consistently provides faithful explanations, thus meeting the needs of the medical field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。