融合皮肤病基础模型与ViT特征,提升皮肤病变分类准确率
Fusion of Foundation and Vision Transformer Model Features for Dermatoscopic Image Classification
- 用冻结的PanDerm特征配合MLP进行非线性探测
- PanDerm+MLP表现接近微调后的Swin Transformer
- 融合两者预测结果可进一步提升性能,适合医学图像分析研究者
从皮肤镜图像中准确分类皮肤病变对皮肤癌的诊断与治疗至关重要。本研究对比了专用于皮肤病的奠基模型PanDerm与两种视觉变换器(ViT base和Swin Transformer V2 base)在皮肤病变分类任务中的表现。使用冻结的PanDerm特征,采用多层感知机(MLP)、XGBoost和TabNet三种分类器进行非线性探测;对于ViT模型,则执行全量微调以优化分类性能。在HAM10000和MSKCC数据集上的实验表明,基于PanDerm的MLP模型表现与微调后的Swin Transformer相当,而融合PanDerm与Swin Transformer的预测结果可带来进一步性能提升。未来工作将探索更多奠基模型、微调策略及先进融合技术。
原文摘要 · Abstract (English)
Accurate classification of skin lesions from dermatoscopic images is essential for diagnosis and treatment of skin cancer. In this study, we investigate the utility of a dermatology-specific foundation model, PanDerm, in comparison with two Vision Transformer (ViT) architectures (ViT base and Swin Transformer V2 base) for the task of skin lesion classification. Using frozen features extracted from PanDerm, we apply non-linear probing with three different classifiers, namely, multi-layer perceptron (MLP), XGBoost, and TabNet. For the ViT-based models, we perform full fine-tuning to optimize classification performance. Our experiments on the HAM10000 and MSKCC datasets demonstrate that the PanDerm-based MLP model performs comparably to the fine-tuned Swin transformer model, while fusion of PanDerm and Swin Transformer predictions leads to further performance improvements. Future work will explore additional foundation models, fine-tuning strategies, and advanced fusion techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。