arXiv:2412.19124cs.CVcs.AI2024-12被引 12

构建医疗影像自监督学习综合评估基准,测试模型鲁棒性与跨域泛化能力。

Evaluating Self-Supervised Learning in Medical Imaging: A Benchmark for Robustness, Generalizability, and Multi-Domain Impact

  • 用 MedMNIST 标准化数据集评估 8 种自监督方法在 11 个医疗数据集上的表现。
  • 在 1%~100% 标注比例下验证模型在小样本场景的性能,发现多域预训练提升泛化性。
  • 首次系统评估模型对分布外样本的检测能力,为临床部署提供可靠性参考。

自监督学习(SSL)在医疗影像领域展现出巨大潜力,可缓解标注数据稀缺问题。然而现有研究多局限于特定数据集或模态,且仅评估模型单一性能指标,难以反映真实医疗环境中的复杂需求。为此,本文以 MedMNIST 数据集集合为标准化基准,系统评估 8 种主流 SSL 方法在 11 个医疗数据集上的表现。研究涵盖同域性能与分布外(OOD)样本检测能力,并考察不同初始化策略、模型架构及多域预训练的影响。通过跨数据集评估和在 1%、10%、100% 标注比例下的性能分析,模拟实际中标签有限的场景。结果表明,多域预训练显著提升模型泛化能力,且部分方法在极低标注率下仍保持良好性能。本研究为医疗 SSL 应用提供了全面的评估框架,助力研究人员与实践者做出更明智选择。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) has emerged as a promising paradigm in medical imaging, addressing the chronic challenge of limited labeled data in healthcare settings. While SSL has shown impressive results, existing studies in the medical domain are often limited in scope, focusing on specific datasets or modalities, or evaluating only isolated aspects of model performance. This fragmented evaluation approach poses a significant challenge, as models deployed in critical medical settings must not only achieve high accuracy but also demonstrate robust performance and generalizability across diverse datasets and varying conditions. To address this gap, we present a comprehensive evaluation of SSL methods within the medical domain, with a particular focus on robustness and generalizability. Using the MedMNIST dataset collection as a standardized benchmark, we evaluate 8 major SSL methods across 11 different medical datasets. Our study provides an in-depth analysis of model performance in both in-domain scenarios and the detection of out-of-distribution (OOD) samples, while exploring the effect of various initialization strategies, model architectures, and multi-domain pre-training. We further assess the generalizability of SSL methods through cross-dataset evaluations and the in-domain performance with varying label proportions (1%, 10%, and 100%) to simulate real-world scenarios with limited supervision. We hope this comprehensive benchmark helps practitioners and researchers make more informed decisions when applying SSL methods to medical applications.

自监督学习医疗影像泛化能力基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。