融合图像与临床数据,提升皮肤病变分类准确率。
Multimodal Skin Lesion Classification with Swin Transformer and Clinical Metadata Fusion
- 用Swin Transformer提取图像特征,结合临床数据进行多模态融合
- 测试准确率达92.55%,少数类表现优异,宏平均F1为91.33%
- 加入校准与不确定性分析,结果更可信,且可解释性强
皮肤病变分类对早期皮肤癌诊断至关重要,但自动化分析仍面临类别不平衡、类间相似及类内差异等挑战。本文提出一种多模态分类框架,将基于Swin Transformer的图像特征与结构化临床元数据融合,通过视觉-上下文联合学习提升诊断性能。在公开数据集上的实验表明,该模型测试准确率达到92.55%,宏平均F1得分为91.33%,在少数类上表现稳健。采用温度缩放作为后处理校准方法,显著降低期望校准误差,提升预测可靠性;同时引入不确定性估计以评估模型置信度。定性可解释性分析显示,模型在推理时聚焦于病灶区域。结果表明,多模态融合结合校准与可解释性分析,为自动皮肤病变分类提供了高效且可信的解决方案。
原文摘要 · Abstract (English)
Skin lesion classification plays an important role in supporting the early diagnosis of skin cancer. However, automated analysis remains challenging due to class imbalance, inter-class similarity, and intra-class variability in dermoscopic images. This paper proposes a multimodal classification framework that combines Swin Transformer-based image features with structured clinical metadata to improve diagnostic performance through integrated visual-context learning. Experiments on a publicly available dataset show that the proposed model achieves a test accuracy of 92.55% and a macro F1-score of 91.33%, with strong performance across minority classes. Temperature scaling is applied as a post-hoc calibration method, resulting in a reduction in expected calibration error and improving prediction reliability, while uncertainty estimation is incorporated to further assess the confidence of model predictions. Qualitative explainability analysis further shows that the model focuses on lesion regions during inference. Therefore, the results demonstrate that multimodal fusion, combined with calibration and interpretability analysis, provides an effective and trustworthy approach for automated skin lesion classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。