arXiv:2604.01987cs.CVcs.LG2026-04被引 5

Curia-2提升医学影像基础模型预训练能力,实现百亿参数级视觉变压器。

Curia-2: Scaling Self-Supervised Learning for Radiology Foundation Models

  • 采用新预训练策略,支持百亿参数视觉变压器在多模态CT/MRI上训练
  • 在2D与3D基准测试中均超越现有视觉基础模型,临床检测任务媲美图文模型
  • 开源权重,推动医学影像自监督学习研究

医学影像的快速增长推动了基础模型(FMs)的发展,以缓解放射科医生日益增长且不可持续的工作负担。尽管近期基础模型在CT和MRI分析中展现出大规模预训练的潜力,但在如何从复杂的放射学影像中高效学习方面仍有优化空间。本文在Curia框架基础上提出Curia-2,显著改进原始预训练策略与表征质量,更精准捕捉放射学数据特性。该方法首次实现多模态CT/MRI基础模型扩展至百亿参数级别的视觉变压器。同时,我们将CuriaBench重构为两个独立评测赛道:面向切片的2D赛道和面向体数据的3D赛道。实验表明,Curia-2在视觉任务上优于所有现有基础模型,在复杂临床任务如病灶检测上表现媲美视觉-语言模型。模型权重将公开,以促进后续研究。

原文摘要 · Abstract (English)

The rapid growth of medical imaging has fueled the development of Foundation Models (FMs) to reduce the growing, unsustainable workload on radiologists. While recent FMs have shown the power of large-scale pre-training to CT and MRI analysis, there remains significant room to optimize how these models learn from complex radiological volumes. Building upon the Curia framework, this work introduces Curia-2, which significantly improves the original pre-training strategy and representation quality to better capture the specificities of radiological data. The proposed methodology enables scaling the architecture up to billion-parameter Vision Transformers, marking a first for multi-modal CT and MRI FMs. Furthermore, we formalize the evaluation of these models by extending and restructuring CuriaBench into two distinct tracks: a 2D track tailored for slice-based vision models and a 3D track for volumetric benchmarking. Our results demonstrate that Curia-2 outperforms all FMs on vision-focused tasks and fairs competitively to vision-language models on clinically complex tasks such as finding detection. Weights will be made publicly available to foster further research.

医学影像基础模型自监督学习视觉变压器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。