arXiv:2505.09329cs.CVcs.AI2025-05被引 1

构建2100万张生物医学图像数据集,验证大模型在医疗图像中的可扩展性。

BioVFM-21M: Benchmarking and Scaling Self-Supervised Vision Foundation Models for Biomedical Image Analysis

  • 基于自监督学习,系统研究模型规模与数据量对医疗图像的影响
  • 在12个医学基准上超越现有最优模型,展现显著性能提升
  • 适合医疗视觉大模型研发者参考其训练策略与数据设计

扩大模型和数据规模在多个任务中展现出显著性能提升。尽管通用领域已有大量关于缩放行为的研究,但医学图像与自然图像存在显著差异。由于缺乏对医学领域缩放行为的深入理解,开发可扩展的医学视觉基础模型的关键因素尚不明确。本文通过自监督学习,系统探索了模型规模、训练算法、数据规模和成像模态对可扩展医学视觉基础模型的影响。为支持大规模预训练,我们引入BioVFM-21M,一个包含2100万张生物医学图像的大型数据集,涵盖多种成像模态和解剖结构。实验表明,扩大规模虽有益处,但效果因任务而异。进一步分析揭示了若干与缩放收益相关的因素。最终,我们提出BioVFM,一个在2100万张生物医学图像上预训练的大规模医学视觉基础模型,在12个医学基准上优于此前的最先进模型。结果表明,虽然扩大规模有助于提升性能,但任务特性、数据多样性、预训练方法和计算效率仍是构建可扩展医学基础模型的关键考量。

原文摘要 · Abstract (English)

Scaling up model and data size have demonstrated impressive performance improvement over a wide range of tasks. Despite extensive studies on scaling behaviors for general-purpose tasks, medical images exhibit substantial differences from natural data. It remains unclear the key factors in developing medical vision foundation models at scale due to the absence of an extensive understanding of scaling behavior in the medical domain. In this paper, we explored the scaling behavior across model sizes, training algorithms, data sizes, and imaging modalities in developing scalable medical vision foundation models by self-supervised learning. To support scalable pretraining, we introduce BioVFM-21M, a large-scale biomedical image dataset encompassing a wide range of biomedical image modalities and anatomies. We observed that scaling up does provide benefits but varies across tasks. Additional analysis reveals several factors correlated with scaling benefits. Finally, we propose BioVFM, a large-scale medical vision foundation model pretrained on 21 million biomedical images, which outperforms the previous state-of-the-art foundation models across 12 medical benchmarks. Our results highlight that while scaling up is beneficial for pursuing better performance, task characteristics, data diversity, pretraining methods, and computational efficiency remain critical considerations for developing scalable medical foundation models.

医学图像自监督学习大模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。