arXiv:2502.02500eess.IVcs.CV2025-02

提出皮肤影像模型评估新标准,解决方法不统一问题。

The Skin Game: Revolutionizing Standards for AI Dermatology Model Comparison

  • 分析现有方法,发现数据处理与报告存在严重不一致
  • DINOv2在三个数据集上分别取得0.85、0.71、0.84的F1分数
  • 强调标准化流程对临床部署的重要性,适合研究者与医生

皮肤疾病图像分类中的深度学习方法虽表现良好,但方法学挑战阻碍了有效评估。本文提出双重贡献:首先系统分析当前皮肤疾病分类研究的方法学实践,揭示数据准备、增强策略和性能报告中的显著不一致;其次通过在HAM10000、DermNet和ISIC Atlas三个基准数据集上使用DINOv2-Large视觉变换器的实验,展示全面的训练与评估框架。分析发现预分割数据增强和基于验证集的报告方式可能高估指标,且缺乏统一标准。实验结果表明DINOv2在三数据集上的宏平均F1分数分别为0.85(HAM10000)、0.71(DermNet)和0.84(ISIC Atlas)。注意力图分析显示模型对典型病例有较强特征识别能力,但在非典型病例和复合图像上存在明显弱点。研究呼吁建立标准化评估协议,并提出涵盖数据准备、系统误差分析及不同图像类型专用方案的开发与部署建议。为促进可复现性,代码已开源。本工作为皮肤影像分类建立了严谨评估基础,推动临床皮肤病学中负责任的AI应用。

原文摘要 · Abstract (English)

Deep Learning approaches in dermatological image classification have shown promising results, yet the field faces significant methodological challenges that impede proper evaluation. This paper presents a dual contribution: first, a systematic analysis of current methodological practices in skin disease classification research, revealing substantial inconsistencies in data preparation, augmentation strategies, and performance reporting; second, a comprehensive training and evaluation framework demonstrated through experiments with the DINOv2-Large vision transformer across three benchmark datasets (HAM10000, DermNet, ISIC Atlas). The analysis identifies concerning patterns, including pre-split data augmentation and validation-based reporting, potentially leading to overestimated metrics, while highlighting the lack of unified methodology standards. The experimental results demonstrate DINOv2's performance in skin disease classification, achieving macro-averaged F1-scores of 0.85 (HAM10000), 0.71 (DermNet), and 0.84 (ISIC Atlas). Attention map analysis reveals critical patterns in the model's decision-making, showing sophisticated feature recognition in typical presentations but significant vulnerabilities with atypical cases and composite images. Our findings highlight the need for standardized evaluation protocols and careful implementation strategies in clinical settings. We propose comprehensive methodological recommendations for model development, evaluation, and clinical deployment, emphasizing rigorous data preparation, systematic error analysis, and specialized protocols for different image types. To promote reproducibility, we provide our implementation code through GitHub. This work establishes a foundation for rigorous evaluation standards in dermatological image classification and provides insights for responsible AI implementation in clinical dermatology.

皮肤影像AI评估标准化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。