用10万例3D医学影像训练通用模型,提升跨模态跨器官泛化能力。
A generalizable 3D framework and model for self-supervised learning in medical imaging
- 基于3DINO框架,利用自监督学习预训练跨器官多模态模型
- 在超10万例3D医学影像上训练,覆盖10种以上器官和多种成像模态
- 在多个下游任务中优于现有方法,适合医疗影像通用建模与微调
当前3D医学影像的自监督学习方法依赖简单预训练任务和器官或模态特定数据集,限制了其泛化性和可扩展性。本文提出3DINO,一种适用于3D数据集的先进自监督学习方法,并基于超过10个器官、包含约10万例3D医学影像的多模态大规模数据集,预训练出3DINO-ViT——一个通用医学影像模型。通过在多个医学影像分割与分类任务上的广泛实验验证,结果表明3DINO-ViT具备跨模态、跨器官的泛化能力,包括分布外任务和数据集,在多数评估指标和不同标注数据规模下均超越现有最先进方法。3DINO框架及3DINO-ViT将公开,以支持3D基础模型研究及各类医学影像应用的微调。
原文摘要 · Abstract (English)
Current self-supervised learning methods for 3D medical imaging rely on simple pretext formulations and organ- or modality-specific datasets, limiting their generalizability and scalability. We present 3DINO, a cutting-edge SSL method adapted to 3D datasets, and use it to pretrain 3DINO-ViT: a general-purpose medical imaging model, on an exceptionally large, multimodal, and multi-organ dataset of ~100,000 3D medical imaging scans from over 10 organs. We validate 3DINO-ViT using extensive experiments on numerous medical imaging segmentation and classification tasks. Our results demonstrate that 3DINO-ViT generalizes across modalities and organs, including out-of-distribution tasks and datasets, outperforming state-of-the-art methods on the majority of evaluation metrics and labeled dataset sizes. Our 3DINO framework and 3DINO-ViT will be made available to enable research on 3D foundation models or further finetuning for a wide range of medical imaging applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。