对比10个基础模型在皮肤病变分级分类中的表现,发现通用模型适合筛查,专用模型更擅长细粒度诊断。
A Hierarchical Benchmark of Foundation Models for Dermatology
- 构建四级分层评估框架,覆盖从40类亚型到二分类恶性病变的多种临床粒度。
- 通用医学模型在恶性肿瘤二分类上达97.52%准确率,但细粒度分类仅65.50%。
- 皮肤病专用模型在40类细分任务中表现最佳(最高69.79%),适合精准诊断场景。
基础模型通过提供强大的特征表示,改变了医学图像分析范式,减少了对大规模特定任务训练的需求。然而,当前皮肤病学基准常将复杂的诊断分类体系简化为扁平的二分类任务(如黑色素瘤与良性痣区分),这掩盖了模型进行细粒度鉴别诊断的能力,而该能力对临床流程整合至关重要。本研究评估了十个基础模型(涵盖通用计算机视觉、通用医学影像及皮肤病专用领域)在层次化皮肤病变分类中的嵌入效果。基于包含40个病变亚类的DERM12345数据集,计算冻结嵌入并采用五折交叉验证训练轻量级适配器模型。引入一种分层评估框架,从四个临床粒度层级评估性能:40个亚类、15个主类、2和4个超级类,以及二分类恶性病变。结果显示存在‘粒度差距’:MedImageInsights在二分类恶性病变上表现最优(加权F1得分为97.52%),但在细粒度40类子分类上下降至65.50%;而MedSigLip(69.79%)和皮肤病专用模型(Derm Foundation与MONET)在40类子分类中表现优异,但整体分类性能低于MedImageInsights。结果表明,通用医学基础模型适用于高层级筛查,而精细化诊断支持系统需依赖专用建模策略。
原文摘要 · Abstract (English)
Foundation models have transformed medical image analysis by providing robust feature representations that reduce the need for large-scale task-specific training. However, current benchmarks in dermatology often reduce the complex diagnostic taxonomy to flat, binary classification tasks, such as distinguishing melanoma from benign nevi. This oversimplification obscures a model's ability to perform fine-grained differential diagnoses, which is critical for clinical workflow integration. This study evaluates the utility of embeddings derived from ten foundation models, spanning general computer vision, general medical imaging, and dermatology-specific domains, for hierarchical skin lesion classification. Using the DERM12345 dataset, which comprises 40 lesion subclasses, we calculated frozen embeddings and trained lightweight adapter models using a five-fold cross-validation. We introduce a hierarchical evaluation framework that assesses performance across four levels of clinical granularity: 40 Subclasses, 15 Main Classes, 2 and 4 Superclasses, and Binary Malignancy. Our results reveal a "granularity gap" in model capabilities: MedImageInsights achieved the strongest overall performance (97.52% weighted F1-Score on Binary Malignancy detection) but declined to 65.50% on fine-grained 40-class subtype classification. Conversely, MedSigLip (69.79%) and dermatology-specific models (Derm Foundation and MONET) excelled at fine-grained 40-class subtype discrimination while achieving lower overall performance than MedImageInsights on broader classification tasks. Our findings suggest that while general medical foundation models are highly effective for high-level screening, specialized modeling strategies are necessary for the granular distinctions required in diagnostic support systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。