用对比学习让皮肤癌模型说清楚判断依据,提升医生信任度。
Explainable Melanoma Diagnosis with Contrastive Learning and LLM-based Report Generation
- 将临床ABC标准映射到视觉特征空间,实现图像与诊断规则对齐。
- 在公开数据集上达到92.79%准确率和0.961 AUC,解释性显著提升。
- 生成结构化文本报告,适合需要可解释性的医疗AI应用。
深度学习在黑色素瘤分类中已达到专家水平,但在临床中因模型不透明而难以被采纳。为此,我们提出跨模态可解释框架CEFM,以对比学习为核心,将黑色素瘤诊断的临床标准——不对称性、边界、颜色(ABC)——通过双投影头映射至视觉变换器嵌入空间,实现临床语义与视觉特征的对齐。对齐后的表示通过自然语言生成转化为结构化文本解释,建立原始图像与临床判断之间的透明关联。在公开数据集上的实验表明,该方法实现了92.79%的准确率和0.961的AUC,且在多个可解释性指标上均有显著提升。定性分析显示,学习到的嵌入空间布局与临床医生应用ABC规则的方式高度一致,有效弥合了高性能分类与临床信任之间的鸿沟。
原文摘要 · Abstract (English)
Deep learning has demonstrated expert-level performance in melanoma classification, positioning it as a powerful tool in clinical dermatology. However, model opacity and the lack of interpretability remain critical barriers to clinical adoption, as clinicians often struggle to trust the decision-making processes of black-box models. To address this gap, we present a Cross-modal Explainable Framework for Melanoma (CEFM) that leverages contrastive learning as the core mechanism for achieving interpretability. Specifically, CEFM maps clinical criteria for melanoma diagnosis-namely Asymmetry, Border, and Color (ABC)-into the Vision Transformer embedding space using dual projection heads, thereby aligning clinical semantics with visual features. The aligned representations are subsequently translated into structured textual explanations via natural language generation, creating a transparent link between raw image data and clinical interpretation. Experiments on public datasets demonstrate 92.79% accuracy and an AUC of 0.961, along with significant improvements across multiple interpretability metrics. Qualitative analyses further show that the spatial arrangement of the learned embeddings aligns with clinicians' application of the ABC rule, effectively bridging the gap between high-performance classification and clinical trust.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。