AI模型在17年老样本上稳定分级,媲美病理医生。
Validation of an AI-based end-to-end model for prostate pathology using long-term archived routine samples
- 用注意力机制的端到端模型分析前列腺活检组织。
- 核心级分级一致率达0.86,17年间性能无下降。
- 适合做病理大数据研究和长期预后分析。
人工智能在前列腺病理诊断中日益重要,但其在长期保存样本中跨制备与保存差异的泛化能力仍不明确。我们评估了GleasonAI——一种基于注意力机制的端到端多实例学习模型,在包含10,366个活检核心、1,028名患者、来自瑞典14个地区的独立验证队列上的表现,样本来自1998-2015年收集的ProMort队列。模型在核心级ISUP分级上总体加权卡帕值达0.86,与多位经验丰富的病理医生相当,且在各地理区域间保持一致。值得注意的是,模型性能在长达17年的采集周期内稳定,对时间相关的档案材料变异具有鲁棒性,而基础模型方法常不具备此特性。探索性分析显示,AI分配的分级组与前列腺癌特异性死亡率存在显著预后梯度。这些结果支持该AI分级模型的泛化能力,并证明病理档案可作为大规模AI开发、验证及回顾性预后研究的重要资源。
原文摘要 · Abstract (English)
Artificial intelligence (AI) is becoming a clinical tool for prostate pathology, but generalization across variations in sample preparation and preservation over prolonged time periods remains poorly understood. We evaluated GleasonAI, an end-to-end attention-based multiple instance learning model, on an independent validation cohort comprising 10,366 biopsy cores from 1,028 patients across 14 Swedish regions, using archival diagnostic specimens from the ProMort cohorts collected between 1998-2015. The model achieved an overall quadratic-weighted kappa of 0.86 for core-level ISUP grading, comparable to several experienced pathologists and consistent across geographic regions. Notably, performance remained stable across the 17-year collection period, demonstrating robustness to time-related variation in archival material, a property not consistently observed with foundation model-based approaches, with exploratory analysis demonstrating a significant prognostic gradient across AI-assigned grade groups for prostate cancer-specific mortality. These findings support the generalizability of the AI grading model and demonstrate the potential of pathology archives as a large-scale resource for AI development, validation, and retrospective prognostic research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。