arXiv:2412.17041cs.CVcs.AI2024-12ICCV被引 25

构建最大3D脑部MRI自监督预训练数据集,统一评估标准

An OpenMind for 3D medical vision self-supervised learning

  • 发布11.4万例3D脑MRI数据集,支持大规模预训练
  • 在统一框架下验证多种方法,部分超越从零开始的nnU-Net基准
  • 开源代码与模型,推动领域可复现与快速迭代

3D医学图像自监督学习(SSL)领域缺乏一致性与标准化。尽管已有多种方法,但由于预训练数据集规模小且不一致、模型架构各异,以及下游评估数据集不同,难以确定当前最佳方法。本文通过三项关键贡献推动该领域发展:首先,公开目前最大规模的公共预训练数据集,包含114,000个3D脑部MRI体积,使所有研究者可基于大规模数据进行预训练;其次,在该数据集上对现有3D自监督学习方法进行了基准测试,涵盖最先进的CNN与Transformer架构,结果显示部分预训练方法性能超越强基线nnU-Net ResEnc-L;最后,公开预训练与微调框架代码,及由此产生的预训练模型,以促进方法快速采用与复现。

原文摘要 · Abstract (English)

The field of self-supervised learning (SSL) for 3D medical images lacks consistency and standardization. While many methods have been developed, it is impossible to identify the current state-of-the-art, due to i) varying and small pretraining datasets, ii) varying architectures, and iii) being evaluated on differing downstream datasets. In this paper, we bring clarity to this field and lay the foundation for further method advancements through three key contributions: We a) publish the largest publicly available pre-training dataset comprising 114k 3D brain MRI volumes, enabling all practitioners to pre-train on a large-scale dataset. We b) benchmark existing 3D self-supervised learning methods on this dataset for a state-of-the-art CNN and Transformer architecture, clarifying the state of 3D SSL pre-training. Among many findings, we show that pre-trained methods can exceed a strong from-scratch nnU-Net ResEnc-L baseline. Lastly, we c) publish the code of our pre-training and fine-tuning frameworks and provide the pre-trained models created during the benchmarking process to facilitate rapid adoption and reproduction.

3D医学影像自监督学习数据集预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。