arXiv:2509.12818cs.CVcs.AI2025-09被引 5

用350万张胸片持续预训练医学视觉模型,发现数据量对不同任务影响显著。

Data Scaling Laws for Radiology Foundation Models

  • 在单机构350万张胸片上持续预训练两个医学模型,固定算力与评估方式。
  • 仅3万张内部数据即可超越开放权重模型,且报告+结构化标签提升性能。
  • 针对病灶的任务适合CLIP类模型,管状结构则更适合DINO类模型。

CLIP和DINOv2等通用视觉编码器在大规模网络数据上预训练后表现出强迁移能力,但医学影像基础模型受限于较小数据集,其数据规模与预训练范式对性能的影响尚不明确。本文系统研究了两种主流编码器——代表CLIP范式的MedImageInsight(MI2)和代表DINOv2范式的RAD-DINO——在单机构350万张胸部X光片上的持续预训练效果,保持计算资源与评估协议一致。评估涵盖分类(放射学发现、导线与导管)、分割(导线与导管)及放射科报告生成任务。不同于以往集中于病灶识别的研究,本工作引入导线与导管任务以平衡偏见,并考察模型对长条状结构连续特征的提取能力。实验表明,MI2在病灶相关任务中表现更优,而RAD-DINO在导管相关任务中更具优势。令人意外的是,使用UniCL联合报告与结构化标签持续预训练MI2能显著提升性能,凸显大规模结构化监督的价值。进一步发现,部分任务仅需3万张领域内样本即可超越开放权重基础模型。这些结果强调了中心特异性持续预训练的潜力,使医疗机构可通过自有数据实现显著性能提升。

原文摘要 · Abstract (English)

Foundation vision encoders such as CLIP and DINOv2, trained on web-scale data, exhibit strong transfer performance across tasks and datasets. However, medical imaging foundation models remain constrained by smaller datasets, limiting our understanding of how data scale and pretraining paradigms affect performance in this setting. In this work, we systematically study continual pretraining of two vision encoders, MedImageInsight (MI2) and RAD-DINO representing the two major encoder paradigms CLIP and DINOv2, on up to 3.5M chest x-rays from a single institution, holding compute and evaluation protocols constant. We evaluate on classification (radiology findings, lines and tubes), segmentation (lines and tubes), and radiology report generation. While prior work has primarily focused on tasks related to radiology findings, we include lines and tubes tasks to counterbalance this bias and evaluate a model's ability to extract features that preserve continuity along elongated structures. Our experiments show that MI2 scales more effectively for finding-related tasks, while RAD-DINO is stronger on tube-related tasks. Surprisingly, continually pretraining MI2 with both reports and structured labels using UniCL improves performance, underscoring the value of structured supervision at scale. We further show that for some tasks, as few as 30k in-domain samples are sufficient to surpass open-weights foundation models. These results highlight the utility of center-specific continual pretraining, enabling medical institutions to derive significant performance gains by utilizing in-domain data.

医学影像数据规模持续预训练模型对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。