arXiv:2411.02372cs.CVcs.LG2024-11ICLR被引 35

用随机合成数据训练3D医学图像模型,实现跨场景泛化。

Learning General-Purpose Biomedical Volume Representations using Randomized Synthesis

  • 通过随机合成多样化数据增强训练样本多样性。
  • 对比学习使模型对成像差异保持稳定,提升泛化能力。
  • 无需真实数据预训练,适用于新任务和新数据集。

当前三维医学基础模型难以泛化,因公开数据集规模小且覆盖范围有限。本文提出一种表示学习方法,在训练阶段主动模拟强域偏移。首先设计数据引擎,生成高度多样化的合成训练样本以支持跨医学场景泛化;随后开发对比学习方法,使单个3D网络在数据引擎模拟的噪声与变化下保持特征稳定,这是实现泛化的重要归纳偏置。该网络提取的特征可作为下游任务的鲁棒表示,其权重也可作为新数据集微调的强初始化。结果在多模态配准与少样本分割上均达到新标准,首次实现无需任何真实图像(预)训练的3D医学视觉模型性能突破。

原文摘要 · Abstract (English)

Current volumetric biomedical foundation models struggle to generalize as public 3D datasets are small and do not cover the broad diversity of medical procedures, conditions, anatomical regions, and imaging protocols. We address this by creating a representation learning method that instead anticipates strong domain shifts at training time itself. We first propose a data engine that synthesizes highly variable training samples that would enable generalization to new biomedical contexts. To then train a single 3D network for any voxel-level task, we develop a contrastive learning method that pretrains the network to be stable against nuisance imaging variation simulated by the data engine, a key inductive bias for generalization. This network's features can be used as robust representations of input images for downstream tasks and its weights provide a strong, dataset-agnostic initialization for finetuning on new datasets. As a result, we set new standards across both multimodality registration and few-shot segmentation, a first for any 3D biomedical vision model, all without (pre-)training on any existing dataset of real images.

3D医学对比学习数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。