arXiv:2607.07219cs.CVcs.AI2026-07综述

梳理放射科视觉基础模型研究现状与挑战

Vision Foundation Models in Radiology: A Scoping Review of Data, Methodology, Evaluation and Clinical Translation

论文配图:Vision Foundation Models in Radiology: A Scoping Review of Data, Methodology, Evaluation and Clinical Translation
图 1 · 摘自论文原文
  • 系统分析67项放射科图像基础模型研究,聚焦数据、方法与评估
  • 多数采用Transformer和自监督预训练,跨中心验证仍不充分
  • 适合医学AI研究者参考,关注临床落地的瓶颈与改进方向

视觉基础模型(VFMs)在放射学影像中日益发展,但其定义、开发与评估仍存在较大异质性。本研究采用PRISMA-ScR框架,对2017年1月至2026年3月间发表的67项同行评审研究进行了系统性综述,这些研究均基于放射科影像数据训练基础模型。研究从三个维度进行映射:数据规模与多样性、架构与预训练可扩展性、下游任务迁移与泛化能力。数据集主要覆盖脑部MRI、胸腹CT和胸部X光,样本量从不足10万到数百万图像不等。以Transformer架构和自监督预训练为主,尤其是掩码图像建模、对比学习与多阶段方法。评估集中于分割与分类任务,而跨中心、跨扫描仪、解剖结构与模态迁移验证报告不一致。与FUTURE-AI原则的对齐程度参差不齐。总体而言,放射科专用基础模型展现出良好迁移潜力,但临床转化受限于数据代表性不足、基准不统一、报告不完整及缺乏面向部署的评估。

原文摘要 · Abstract (English)

Vision foundation models (VFMs) are increasingly being developed for radiological imaging, yet their definition, development and evaluation remain heterogeneous. We conducted a PRISMAScR scoping review of peer-reviewed studies published between January 2017 and March 2026 describing foundation models trained exclusively on radiological imaging data. Sixty-seven studies were included and mapped across three pillars: data scale and heterogeneity, architectural and pretraining scalability, and downstream transferability and generalization. Datasets primarily covered brain MRI, thoracoabdominal CT, and chest X-ray, ranging from fewer than 100,000 samples to multi-million-image cohorts. Transformer-based architectures and self-supervised pretraining predominated, particularly masked image modeling, contrastive learning and multi-stage approaches. Evaluation focused mainly on segmentation and classification, whereas cross-center, cross-scanner, anatomical and modality-shift validation was inconsistently reported. Alignment with FUTURE-AI principles was uneven. Overall, radiology-specific VFMs show promising transferability, but clinical translation remains constrained by limited data representativeness, heterogeneous benchmarks, incomplete reporting and insufficient deployment-oriented evaluation.

视觉基础模型医学影像临床转化综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。