arXiv:2511.17803cs.CVcs.AI2025-11被引 26

Pillar-0突破放射影像基础模型瓶颈,实现更高精度与临床可用性。

Pillar-0: A New Frontier for Radiology Foundation Models

  • 基于42,990例腹部盆腔CT等数据预训练,全体积处理避免信息丢失
  • 在14,230例腹部CT上达86.4 AUROC,超越现有模型7.8-15.8点
  • 支持真实临床任务如肺癌风险预测,仅用1/20数据即达超优性能

放射影像在现代医学中至关重要,但影像数量增长远超人力。现有医学模型多将三维CT/MRI降维为低质量二维切片,丢失关键灰度对比信息,且缺乏符合临床实践的评估框架。本文提出Pillar-0,一个在42,990例腹部盆腔CT、86,411例胸部CT、14,348例头部CT和11,543例乳腺MRI上预训练的放射影像基础模型,并引入RATE框架,利用大语言模型以近乎完美准确率提取366种放射学发现的结构化标签。在包含14,230例腹部盆腔CT、10,646例胸部CT、4,906例头部CT和1,585例乳腺MRI的内部测试集上,Pillar-0实现86.4、88.0、90.1和82.9的平均AUROC,相比MedGemma、MedImageInsight、Lingshu和Merlin提升7.8–15.8个点,在366项任务中胜出319项(87.2%)。在斯坦福腹部CT外部验证中亦优于所有基线,包括Merlin(82.2 vs 80.6 AUROC)。Pillar-0还拓展至长时程肺癌风险预测,相较SIBYL在NLST数据集上提升3.0 C-index;在MGH和CGMH数据集上分别提升5.9和1.9。在脑出血检测中,仅需基线1/20数据即达>95 AUROC。Pillar-0与RATE共同构建开放、临床严谨的基础平台,突破计算、数据与评估限制,推动此前不可行的应用落地。

原文摘要 · Abstract (English)

Radiology plays an integral role in modern medicine, yet rising imaging volumes have far outpaced workforce growth. Foundation models offer a path toward assisting with the full spectrum of radiology tasks, but existing medical models remain limited: they process volumetric CT and MRI as low-fidelity 2D slices, discard critical grayscale contrast information, and lack evaluation frameworks that reflect real clinical practice. We introduce Pillar-0, a radiology foundation model pretrained on 42,990 abdomen-pelvis CTs, 86,411 chest CTs, 14,348 head CTs, and 11,543 breast MRIs from a large academic center, together with RATE, a scalable framework that extracts structured labels for 366 radiologic findings with near-perfect accuracy using LLMs. Across internal test sets of 14,230 abdomen-pelvis CTs, 10,646 chest CTs, 4,906 head CTs, and 1,585 breast MRIs, Pillar-0 establishes a new performance frontier, achieving mean AUROCs of 86.4, 88.0, 90.1, and 82.9, outperforming MedGemma (Google), MedImageInsight (Microsoft), Lingshu (Alibaba), and Merlin (Stanford) by 7.8-15.8 AUROC points and ranking best in 87.2\% (319/366) tasks. Pillar-0 similarly outperforms all baselines in an external validation on the Stanford Abdominal CT dataset, including Merlin (82.2 vs 80.6 AUROC). Pillar-0 extends to tasks beyond its pretraining, such as long-horizon lung cancer risk prediction, where it improves upon the state-of-the-art Sybil by 3.0 C-index points on NLST, and generalizes with gains of 5.9 (MGH) and 1.9 (CGMH). In brain hemorrhage detection, Pillar-0 obtained a >95 AUROC when using only 1/20th of the data of the next most sample efficient baseline. Pillar-0 and RATE together provide an open, clinically rigorous foundation for building high-performance radiology systems, enabling applications that were previously infeasible due to computational, data, and evaluation constraints.

放射影像基础模型医疗AI临床应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。