arXiv:2506.19552cs.CVcs.AI2025-06中稿 · MICCAI 2025被引 14

用通用方法在胎儿超声数据上训练出领先模型,证明无需创新算法也能高效建模。

General Methods Make Great Domain-specific Foundation Models: A Case-study on Fetal Ultrasound

  • 选用成熟视觉模型DINOv2,在200万张胎儿超声图上预训练
  • 在3个跨国数据集上达成最新性能,覆盖分类、分割与少样本任务
  • 无需调参或改造方法,通用技术即可打造优质医疗专用模型

拥有大规模未标注医学数据后,研究者面临两个问题:是否应针对该医学数据定制预训练基础模型,还是使用现有通用模型进行迁移学习?若选择自定义预训练,是否需要新方法?本文通过一项案例研究回答这些问题:我们在包含200万张图像的区域性胎儿超声数据集上,采用已成熟的DINOv2方法进行预训练,实现了在三个不同国家的胎儿超声数据集上的最先进性能,涵盖分类、分割和少样本任务。我们对比了自然图像、超声图像预训练模型及监督基线模型。结果表明:(i) 即使使用较小模型在较少数据上预训练,定制化数据预训练也值得投入,因为自然图像中的缩放规律无法直接转化为超声表现;(ii) 经良好调优的计算机视觉通用方法已足以支持特定医学领域的基础模型构建,无需复杂参数调整或方法修改。因此,当计算资源有限时,应避免过度追求方法创新。

原文摘要 · Abstract (English)

With access to large-scale, unlabeled medical datasets, researchers are confronted with two questions: Should they attempt to pretrain a custom foundation model on this medical data, or use transfer-learning from an existing generalist model? And, if a custom model is pretrained, are novel methods required? In this paper we explore these questions by conducting a case-study, in which we train a foundation model on a large regional fetal ultrasound dataset of 2M images. By selecting the well-established DINOv2 method for pretraining, we achieve state-of-the-art results on three fetal ultrasound datasets, covering data from different countries, classification, segmentation, and few-shot tasks. We compare against a series of models pretrained on natural images, ultrasound images, and supervised baselines. Our results demonstrate two key insights: (i) Pretraining on custom data is worth it, even if smaller models are trained on less data, as scaling in natural image pretraining does not translate to ultrasound performance. (ii) Well-tuned methods from computer vision are making it feasible to train custom foundation models for a given medical domain, requiring no hyperparameter tuning and little methodological adaptation. Given these findings, we argue that a bias towards methodological innovation should be avoided when developing domain specific foundation models under common computational resource constraints.

医学影像基础模型超声DINOv2

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。