arXiv:2602.12317q-bio.QMcs.AI2026-02

用随机合成与解耦让医学影像模型仅靠生成数据就能高效预训练。

Free Lunch in Medical Image Foundation Model Pre-training via Randomized Synthesis and Disentanglement

  • 通过随机高斯分布模拟解剖结构和外观变化,增强模型对真实特征的泛化能力。
  • 在120万3D体积和960万2D图像上预训练,覆盖48个数据集、56项任务,表现优于从零开始训练。
  • 无需真实标注数据,实现隐私保护、可扩展且临床通用的模型构建,适合医疗AI研发者使用。

医学影像基础模型(MIFMs)在多种临床任务中展现出巨大潜力,但其发展受限于大规模标注数据的稀缺性、异质性及高昂成本。本文提出RaSD(随机合成与解耦)框架,完全基于合成数据进行MIFMs的预训练。通过采用随机高斯分布建模解剖结构与外观变化,RaSD引入多尺度的结构与外观扰动,迫使模型依赖于不变且任务相关的解剖线索,而非数据集特有的纹理,从而实现鲁棒且可迁移的表征学习。我们在120万3D体积和960万2D图像上预训练了RaSD模型,并在6种成像模态、48个数据集和56项下游任务上进行了广泛评估。在所有评估任务中,RaSD均优于从零开始训练的模型,在17项任务上达到最佳性能,多数情况下与在大规模真实数据上预训练的模型相当。结果表明,仅靠合成数据即可驱动稳健的表征学习。本研究确立了医学AI范式转变:合成数据可作为可扩展、隐私保护且临床泛化的基础模型的‘免费午餐’。

原文摘要 · Abstract (English)

Medical image foundation models (MIFMs) have demonstrated remarkable potential for a wide range of clinical tasks, yet their development is constrained by the scarcity, heterogeneity, and high cost of large-scale annotated datasets. Here, we propose RaSD (Randomized Synthesis and Disentanglement), a scalable framework for pre-training MIFMs entirely on synthetic data. By modeling anatomical structures and appearance variations with randomized Gaussian distributions, RaSD exposes models to sufficient multi-scale structural and appearance perturbations, forcing them to rely on invariant and task-relevant anatomical cues rather than dataset-specific textures, thereby enabling robust and transferable representation learning. We pre-trained RaSD on 1.2 million 3D volumes and 9.6 million 2D images, and extensively evaluated the resulting models across 6 imaging modalities, 48 datasets, and 56 downstream tasks. Across all evaluated downstream tasks, RaSD consistently outperforms training-from-scratch models, achieves the best performance on 17 tasks, and remains comparable to models pre-trained on large real datasets in most others. These results demonstrate that the capacity of synthetic data alone to drive robust representation learning. Our findings establish a paradigm shift in medical AI, demonstrating that synthetic data can serve as a "free lunch" for scalable, privacy-preserving, and clinically generalizable foundation models.

医学影像生成模型预训练合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。