用新3D医学数据集训练模型,显著提升医疗影像迁移效果。
How Well Do Supervised 3D Models Transfer to Medical Imaging Tasks?
- 构建9262个3D CT数据集AbdomenAtlas 1.1,含25个解剖结构标注。
- 仅用21个标注病例训练的模型,性能媲美用5050个未标注病例训练的模型。
- 适合关注医学影像、3D模型预训练的研究者和开发者。
预训练-微调范式在迁移学习中广泛应用。例如,在ImageNet上预训练后微调至PASCAL的任务,可显著优于从头训练。然而,ImageNet为2D图像分类数据集,其特征难以直接迁移到3D医学影像分割等任务,导致性能下降。当前缺乏类似ImageNet规模的高质量3D医学标注数据集。为此,本文提出两项贡献:首先构建了AbdomenAtlas 1.1,包含9,262个三维CT体积,对25个解剖结构进行体素级标注,并为7种肿瘤类型提供伪标注;其次开发了基于该数据集预训练的模型套件。初步分析表明,仅用21个标注病例(672张掩码)和40 GPU小时训练的模型,其迁移能力已接近使用5,050个未标注病例和1,152 GPU小时训练的模型。更重要的是,监督模型的迁移能力随标注数据量增加而持续提升,性能显著优于现有预训练模型,无论其预训练方法或数据来源如何。本研究旨在推动更大规模3D医学数据集的构建与监督预训练模型的发布。
原文摘要 · Abstract (English)
The pre-training and fine-tuning paradigm has become prominent in transfer learning. For example, if the model is pre-trained on ImageNet and then fine-tuned to PASCAL, it can significantly outperform that trained on PASCAL from scratch. While ImageNet pre-training has shown enormous success, it is formed in 2D, and the learned features are for classification tasks; when transferring to more diverse tasks, like 3D image segmentation, its performance is inevitably compromised due to the deviation from the original ImageNet context. A significant challenge lies in the lack of large, annotated 3D datasets rivaling the scale of ImageNet for model pre-training. To overcome this challenge, we make two contributions. Firstly, we construct AbdomenAtlas 1.1 that comprises 9,262 three-dimensional computed tomography (CT) volumes with high-quality, per-voxel annotations of 25 anatomical structures and pseudo annotations of seven tumor types. Secondly, we develop a suite of models that are pre-trained on our AbdomenAtlas 1.1 for transfer learning. Our preliminary analyses indicate that the model trained only with 21 CT volumes, 672 masks, and 40 GPU hours has a transfer learning ability similar to the model trained with 5,050 (unlabeled) CT volumes and 1,152 GPU hours. More importantly, the transfer learning ability of supervised models can further scale up with larger annotated datasets, achieving significantly better performance than preexisting pre-trained models, irrespective of their pre-training methodologies or data sources. We hope this study can facilitate collective efforts in constructing larger 3D medical datasets and more releases of supervised pre-trained models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。