提出可调敏感度的级联分类框架,提升皮肤癌诊断临床适用性。
Cascade Classification of Dermoscopic Images of Skin Neoplasms with Controllable Sensitivity and External Clinical Validation

- 采用两级级联结构:先二分类筛查恶性,再细分三类肿瘤。
- 在俄临床数据集上,模型敏感度降至0.53-0.67,显示显著泛化差距。
- 可调节阈值实现敏感度控制,更贴近真实临床诊断逻辑。
目的:对比深度学习架构与分类方案在皮肤肿瘤皮肤镜图像中的表现,并评估其从公开国际数据集向俄罗斯临床数据集迁移的泛化能力。方法:比较四种架构(ViT-B/16、Swin-S、ConvNeXt-S、EfficientNetV2-S)在三种分类方案下的性能:二分类(恶性/良性)、单阶段四分类(良性、MEL、SCC、BCC)和两级级联(二分类初筛后三分类区分MEL/SCC/BCC)。所有模型使用ImageNet预训练权重,统一数据增强策略,在整合的ISIC Archive数据上训练,并在内部保留样本及两个临床数据集(Melanoscope AI移动系统;谢切诺夫大学)上评估。结果:内部二分类阶段ROC-AUC为0.952–0.966;在谢切诺夫大学数据集上,AUC降至0.797–0.893,敏感度降至0.53–0.67,ECE从0.02升至0.27–0.39,出现恶性低估现象,量化了排序与校准上的泛化差距。配对检验确认在临床数据上,ViT-B/16在二分类阶段存在显著缺陷(p<0.05);在分化阶段无架构优势。级联结构在多数架构上提升宏F1,但仅对ViT-B/16有显著提升,通过恢复被误判为良性类别的恶性病变实现。在ISIC MILK10k上,直接11类分类平均类别敏感度为0.525。结论:可调阈值提供标准单阶段(argmax)分类无法实现的敏感度控制,更贴合临床鉴别诊断逻辑。持续存在的泛化差距要求部署前进行外部临床验证与再校准。
原文摘要 · Abstract (English)
Purpose. To compare deep learning architectures and classification schemes for dermoscopic images of skin neoplasms and assess their generalization on transfer from open international datasets to independent clinical datasets of Russian practice. Methods. Four architectures (ViT-B/16, Swin-S, ConvNeXt-S, EfficientNetV2-S) were compared in three schemes: binary (malignant/benign), single-stage four-class (benign, MEL, SCC, BCC), and a two-stage cascade (binary triage, then three-class differentiation MEL/SCC/BCC). All models used ImageNet-pretrained weights and a single augmentation protocol on aggregated open ISIC Archive data, and were evaluated on an internal held-out sample and two clinical datasets (Melanoscope AI mobile system; Sechenov University). Results. Internally the binary stage attains ROC-AUC 0.952-0.966; on Sechenov University it drops to 0.797-0.893, sensitivity to 0.53-0.67, and ECE rises from 0.02 to 0.27-0.39 with underestimation of malignancy, quantifying a generalization gap in ranking and calibration. Paired tests confirm one inter-architecture result on clinical data: the deficit of ViT-B/16 at the binary stage (p<0.05); at the differentiation stage no architecture has a proven advantage. The cascade raises macro F1 over single-stage four-class classification for most architectures, but significantly only for ViT-B/16, by recovering malignant lesions assigned to the dominant benign class. On ISIC MILK10k, direct 11-class classification yields mean-class sensitivity 0.525. Conclusion. A tunable triage threshold gives sensitivity control not attainable in standard single-stage (argmax) classification and better reproduces clinical differential-diagnosis logic. The persistent generalization gap mandates external clinical validation and recalibration before deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。