arXiv:2501.09001eess.IVcs.CV2025-01被引 56

CT-FM是首个专为CT影像设计的大规模基础模型,可通用处理多种放射科任务。

Vision Foundation Models for Computed Tomography

  • 基于14.8万张CT扫描数据,用无标签对比学习预训练3D模型
  • 在分割、分诊、检索等四项任务上均超越现有最佳模型
  • 能自动识别解剖结构相似性,且结果可解释性强,适合临床部署

基础模型(FMs)在放射学中展现出变革潜力,可跨模态完成复杂任务。本文开发了专为放射科任务设计的大型3D图像预训练模型CT-FM,使用来自Imaging Data Commons的14.8万张计算机断层扫描(CT)图像,通过无标签对比学习进行预训练。我们在四个任务类别上评估了CT-FM:全身与肿瘤分割、头颅CT分诊、医学图像检索和语义理解,表现优于当前最优模型。除量化指标提升外,CT-FM还展现出对解剖区域的自然聚类能力,并能识别跨扫描的相似解剖与结构概念。此外,其在重复测试设置下保持稳健性,嵌入向量关联的显著区域合理。本研究证明了大规模医学影像基础模型的价值,通过开源模型权重、代码和数据,旨在推动更灵活、可靠且可解释的放射科AI解决方案。

原文摘要 · Abstract (English)

Foundation models (FMs) have shown transformative potential in radiology by performing diverse, complex tasks across imaging modalities. Here, we developed CT-FM, a large-scale 3D image-based pre-trained model designed explicitly for various radiological tasks. CT-FM was pre-trained using 148,000 computed tomography (CT) scans from the Imaging Data Commons through label-agnostic contrastive learning. We evaluated CT-FM across four categories of tasks, namely, whole-body and tumor segmentation, head CT triage, medical image retrieval, and semantic understanding, showing superior performance against state-of-the-art models. Beyond quantitative success, CT-FM demonstrated the ability to cluster regions anatomically and identify similar anatomical and structural concepts across scans. Furthermore, it remained robust across test-retest settings and indicated reasonable salient regions attached to its embeddings. This study demonstrates the value of large-scale medical imaging foundation models and by open-sourcing the model weights, code, and data, aims to support more adaptable, reliable, and interpretable AI solutions in radiology.

CT影像基础模型医学影像3D预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。