arXiv:2501.06887cs.CVcs.AI2025-01中稿 · 2025 IEEE/CVF Wint…被引 2

用视觉语言模型让皮肤癌诊断更透明,看清AI判断依据

MedGrad E-CLIP: Enhancing Trust and Transparency in AI-Driven Skin Lesion Diagnosis

  • 基于CLIP模型融合图像与诊断术语,学习医学特征关联
  • 引入加权熵机制突出关键病变区域,提升解释性
  • 适合医疗AI可信度研究者和临床辅助诊断系统开发者

随着深度学习在医学数据中的广泛应用,确保决策过程的透明与可信至关重要。在皮肤癌诊断中,尽管病变检测与分类的准确率不断提升,但模型的黑箱特性使医生难以理解其判断逻辑,影响信任度。本研究采用在不同皮肤病变数据集上训练的CLIP(对比语言-图像预训练)模型,捕捉视觉特征与诊断标准术语之间的语义关联。为增强可解释性,提出MedGrad E-CLIP方法,基于梯度的E-CLIP框架,引入针对复杂医学影像设计的加权熵机制,突出与特定诊断描述相关的图像关键区域。所构建的集成流程不仅能通过匹配诊断描述实现病变分类,还为医学数据提供了额外的可解释层。通过可视化图像特征与诊断标准之间的联系,该方法展示了先进视觉-语言模型在医学图像分析中的潜力,显著提升了AI诊断系统的透明性、鲁棒性与临床信任度。

原文摘要 · Abstract (English)

As deep learning models gain attraction in medical data, ensuring transparent and trustworthy decision-making is essential. In skin cancer diagnosis, while advancements in lesion detection and classification have improved accuracy, the black-box nature of these methods poses challenges in understanding their decision processes, leading to trust issues among physicians. This study leverages the CLIP (Contrastive Language-Image Pretraining) model, trained on different skin lesion datasets, to capture meaningful relationships between visual features and diagnostic criteria terms. To further enhance transparency, we propose a method called MedGrad E-CLIP, which builds on gradient-based E-CLIP by incorporating a weighted entropy mechanism designed for complex medical imaging like skin lesions. This approach highlights critical image regions linked to specific diagnostic descriptions. The developed integrated pipeline not only classifies skin lesions by matching corresponding descriptions but also adds an essential layer of explainability developed especially for medical data. By visually explaining how different features in an image relates to diagnostic criteria, this approach demonstrates the potential of advanced vision-language models in medical image analysis, ultimately improving transparency, robustness, and trust in AI-driven diagnostic systems.

医学AI可解释性视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。