arXiv:2410.09890cs.CVcs.AI2024-10TPAMI被引 50

利用器官间几何关系自监督预训练,提升3D医学图像分析性能。

Large-Scale 3D Medical Image Pre-training with Geometric Context Priors

  • 通过对比不同区域切片的位置关系进行自监督学习。
  • 在48个任务上实现显著优于基线的性能,最高提升12.7%。
  • 适合医疗影像研究者、模型开发者使用,尤其关注少标注场景。

标注数据稀缺是医学图像分析的主要挑战。大规模预训练因其能利用海量数据、大模型和先进训练技术,成为高效的标签节约解决方案。然而,其在医学图像领域的研究仍不充分。核心难点在于如何利用大量无标注数据并学习高层语义。我们观察到3D医学图像中存在一致的几何上下文,即不同器官间的空间关系具有稳定性,这为学习一致表征提供了新思路。受此启发,我们提出简单而有效的体积对比(VoCo)框架,利用几何上下文先验实现自监督学习。给定输入体数据,从不同区域提取基础切片构建正负样本对,再通过对比随机切片与基础切片的相似性来预测其上下文位置。该方法将内在几何结构编码至模型表示中,无需标注即可促进高层语义学习。具体而言:(1) 构建了目前最大的医学预训练数据集PreCT-160K;(2) 研究缩放规律并提出针对不同医学任务的模型尺寸适配指南;(3) 建立涵盖48项任务的基准测试。大量实验表明VoCo具有明显优势。

原文摘要 · Abstract (English)

The scarcity of annotations poses a significant challenge in medical image analysis. Large-scale pre-training has emerged as a promising label-efficient solution, owing to the utilization of large-scale data, large models, and advanced pre-training techniques. However, its development in medical images remains underexplored. The primary challenge lies in harnessing large-scale unlabeled data and learning high-level semantics without annotations. We observe that 3D medical images exhibit consistent geometric context, i.e., consistent geometric relations between different organs, which leads to a promising way for learning consistent representations. Motivated by this, we introduce a simple-yet-effective Volume Contrast (VoCo) framework to leverage geometric context priors for self-supervision. Given an input volume, we extract base crops from different regions to construct positive and negative pairs for contrastive learning. Then we predict the contextual position of a random crop by contrasting its similarity to the base crops. In this way, VoCo encodes the inherent geometric context into model representations, facilitating high-level semantic learning without annotations. Specifically, we (1) introduce the largest medical pre-training dataset PreCT-160K; (2) investigate scaling laws and propose guidelines for tailoring different model sizes to various medical tasks; (3) build a benchmark encompassing 48 medical tasks. Extensive experiments highlight the superiority of VoCo. Codes at https://github.com/Luffy03/Large-Scale-Medical.

医学图像自监督3D预训练几何先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。