arXiv:2605.25878eess.IVcs.CV2026-05

首个临床验证的肺病理通用模型,可全流程辅助诊断与决策。

A Clinically Validated Foundation Model for Comprehensive Lung Pathology Interpretation

论文配图:A Clinically Validated Foundation Model for Comprehensive Lung Pathology Interpretation
图 1 · 摘自论文原文
  • 基于4万张肺组织切片预训练,覆盖32项临床任务。
  • 前瞻性研究中平均AUC达92.3%,可减少68.8%复审负担。
  • 提升病理医生诊断准确率与效率,支持真实世界应用。

病理评估指导肺癌诊断、治疗选择和预后判断,但现有CPath方法依赖针对单一任务的模型。尽管泛癌基础模型具备泛化能力,却缺乏亚专科深度,且未在临床流程中进行前瞻性验证。我们提出PulmoFoundation,一个跨中心、前瞻性验证、经随机对照试验(RCT)评估的肺病理全面分析基础模型,覆盖术前、术中及术后各阶段。该模型基于Virchow2,通过约4万张诊断性H&E染色全幻灯片(WSI)进行亚专科特异性预训练,在约2.6万张WSI上系统评估了32个临床相关任务。除精准预测分子标志物与患者生存外,其在活检、冰冻切片及手术切除片的核心诊断任务中均达到临床级表现。在注册的1357例患者前瞻性研究中,模型在11个任务上平均AUC为92.3%。使用预设分诊阈值,可使68.8%活检与83.0%冰冻切片无需二次审核,同时推迟44.5%免疫组化染色需求,阳性预测值分别为1.000、0.991和0.966。此外,我们开展交叉随机对照试验,8名病理科医生参与,结果显示AI辅助下诊断准确率从83.2%提升至91.7%(5264个案例对)。诊断时间平均缩短18.3%,诊断信心提升9.0%,组间一致性由中等(kappa=0.55)提升至显著(kappa=0.76)。上述评估共同支持PulmoFoundation作为已临床验证的肺病理决策支持系统。

原文摘要 · Abstract (English)

Pathological assessment guides lung cancer diagnosis, treatment selection, and prognostic evaluation, yet current CPath approaches rely on task-specific models for isolated objectives. Although pan-cancer foundation models offer versatility, they lack subspecialty-level depth and have not been evaluated across clinical workflows or prospectively validated in real-world settings. We introduce PulmoFoundation, a multi-center, prospectively validated, randomized controlled trial (RCT)-evaluated foundation model for comprehensive lung pathology assessment across pre-operative, intra-operative, and post-operative care. Built upon Virchow2 via subspecialty-specific pretraining using ~40,000 diagnostic H&E-stained whole-slide images (WSIs), PulmoFoundation was systematically evaluated on ~26,000 WSIs across 32 clinically relevant tasks. In addition to accurately predicting molecular markers and patient survival, our model achieves clinical-grade performance in core diagnostic tasks across biopsy, frozen section, and surgical resection slides. In a registered prospective study of 1,357 patients across 11 diagnostic tasks, our model achieved an average AUC of 92.3%. Using pre-specified triage thresholds, PulmoFoundation could reduce additional second-review burden for 68.8% of biopsies and 83.0% of frozen sections, and defer 44.5% of IHC stain orders, with PPVs of 1.000, 0.991, and 0.966. Beyond prospective validation, we conducted a crossover RCT with eight pathologists, in which AI assistance improved diagnostic accuracy across 5,264 case-reader pairs (91.7% w/ AI vs. 83.2% w/o AI). AI assistance also reduced median diagnostic time by 18.3%, increased diagnostic confidence by 9.0%, and improved inter-rater agreement from moderate (kappa = 0.55) to substantial (kappa = 0.76). Together, these evaluations support PulmoFoundation as a clinically validated decision-support system for lung pathology.

病理分析AI医疗临床验证基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。