构建标准化病理模型评估框架,提升癌症生物标志物研究可复现性。
GOLDMARK: Governed Outcome-Linked Diagnostic Model Assessment Reference Kit
- 基于TCGA和MSKCC数据集,提供结构化中间结果与评测标准
- 33项任务平均AUC达0.689(TCGA)和0.630(MSKCC)
- 聚焦高表现任务,跨机构稳定性强,适合临床级模型对比
计算生物标志物(CBs)是利用人工智能从H&E全切片图像中提取的组织病理学模式,用于预测治疗反应或预后。近期,基于病理基础模型(PFMs)的滑块级多实例学习(MIL)已成为CB开发的标准基线。尽管性能提升显著,计算病理学仍缺乏标准化的中间数据格式、溯源追踪、检查点规范和可复现的评估指标,制约其临床部署。本文提出GOLDMARK(https://artificialintelligencepathology.org),一个基于精选TCGA队列、带有临床可行动态OncoKB 1-3级生物标志物标签的标准化基准框架。GOLDMARK发布结构化中间表示,包括切片坐标图、来自通用PFMs的每滑块特征嵌入、质量控制元数据、预定义患者级划分、训练好的滑块级模型及评估输出。模型在TCGA上训练,在独立的MSKCC队列上进行双向测试。在33个肿瘤-生物标志物任务中,平均AUROC分别为0.689(TCGA)和0.630(MSKCC)。限制于8个最高性能任务时,平均AUROC分别达到0.831和0.801。这些任务对应已知的形态-基因组关联(如LGG IDH1、COAD MSI/BRAF、THCA BRAF/NRAS、BLCA FGFR3、UCEC PTEN),表现出最强的跨站点一致性。不同通用编码器间的差异较小,远低于任务特异性变异。GOLDMARK为计算病理学建立共享实验基底,支持跨数据集与模型的可复现基准测试。
原文摘要 · Abstract (English)
Computational biomarkers (CBs) are histopathology-derived patterns extracted from hematoxylin-eosin (H&E) whole-slide images (WSIs) using artificial intelligence (AI) to predict therapeutic response or prognosis. Recently, slide-level multiple-instance learning (MIL) with pathology foundation models (PFMs) has become the standard baseline for CB development. While these methods have improved predictive performance, computational pathology lacks standardized intermediate data formats, provenance tracking, checkpointing conventions, and reproducible evaluation metrics required for clinical-grade deployment. We introduce GOLDMARK (https://artificialintelligencepathology.org), a standardized benchmarking framework built on a curated TCGA cohort with clinically actionable OncoKB level 1-3 biomarker labels. GOLDMARK releases structured intermediate representations, including tile coordinate maps, per-slide feature embeddings from canonical PFMs, quality-control metadata, predefined patient-level splits, trained slide-level models, and evaluation outputs. Models are trained on TCGA and evaluated on an independent MSKCC cohort with reciprocal testing. Across 33 tumor-biomarker tasks, mean AUROC was 0.689 (TCGA) and 0.630 (MSKCC). Restricting to the eight highest-performing tasks yielded mean AUROCs of 0.831 and 0.801, respectively. These tasks correspond to established morphologic-genomic associations (e.g., LGG IDH1, COAD MSI/BRAF, THCA BRAF/NRAS, BLCA FGFR3, UCEC PTEN) and showed the most stable cross-site performance. Differences between canonical encoders were modest relative to task-specific variability. GOLDMARK establishes a shared experimental substrate for computational pathology, enabling reproducible benchmarking and direct comparison of methods across datasets and models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。