针对前列腺癌分级中AI模型的鲁棒性问题,构建了专用评测基准。
PANDA-PLUS-Bench: A Clinical Benchmark for Evaluating Robustness of AI Foundation Models in Prostate Cancer Diagnosis
- 基于专家标注的活检切片,设计多分辨率、多增强条件的评测数据集
- 发现所有模型在跨切片识别上准确率下降19.9至26.9个百分点
- 专用于前列腺组织训练的模型在生物特征捕捉和抗干扰能力上表现最优
人工智能基础模型在前列腺癌格里森分级中应用日益广泛,其中GP3与GP4的区分直接影响治疗决策。然而,这些模型可能通过学习标本特异性伪影而非可泛化的生物学特征来获得高验证准确率,限制了其实际临床价值。本文提出PANDA-PLUS-Bench,一个源自专家标注前列腺活检切片的精选基准数据集,专门用于量化这一失效模式。该基准包含九例不同患者的全切片图像,涵盖多样化的格里森模式,从中提取了512x512与224x224像素分辨率的非重叠组织块,在八种增强条件下进行测试。我们评估了七种基础模型在分离生物信号与切片级混杂因素方面的能力。结果表明模型间鲁棒性差异显著:Virchow2在大规模模型中滑片编码最低(81.0%),但跨滑片准确率仅47.2%,为第二低;而专为前列腺组织训练的HistoEncoder展现出最高跨滑片准确率(59.7%)和最强滑片编码能力(90.3%),表明特定领域训练有助于提升对生物特征的捕捉与对滑片特异性噪声的抵抗能力。所有模型均表现出明显的组内与跨滑片准确率差距,幅度介于19.9至26.9个百分点之间。我们开源了Google Colab笔记本,支持研究者使用标准化指标评估其他基础模型。PANDA-PLUS-Bench填补了基础模型评估中的关键空白,为格里森分级这一临床重要场景提供专用鲁棒性评估资源。
原文摘要 · Abstract (English)
Artificial intelligence foundation models are increasingly deployed for prostate cancer Gleason grading, where GP3/GP4 distinction directly impacts treatment decisions. However, these models may achieve high validation accuracy by learning specimen-specific artifacts rather than generalizable biological features, limiting real-world clinical utility. We introduce PANDA-PLUS-Bench, a curated benchmark dataset derived from expert-annotated prostate biopsies designed specifically to quantify this failure mode. The benchmark comprises nine carefully selected whole slide images from nine unique patients containing diverse Gleason patterns, with non-overlapping tissue patches extracted at both 512x512 and 224x224 pixel resolutions across eight augmentation conditions. Using this benchmark, we evaluate seven foundation models on their ability to separate biological signal from slide-level confounders. Our results reveal substantial variation in robustness across models: Virchow2 achieved the lowest slide-level encoding among large-scale models (81.0%) yet exhibited the second-lowest cross-slide accuracy (47.2%). HistoEncoder, trained specifically on prostate tissue, demonstrated the highest cross-slide accuracy (59.7%) and the strongest slide-level encoding (90.3%), suggesting tissue-specific training may enhance both biological feature capture and slide-specific signatures. All models exhibited measurable within-slide vs. cross-slide accuracy gaps, though the magnitude varied from 19.9 percentage points to 26.9 percentage points. We provide an open-source Google Colab notebook enabling researchers to evaluate additional foundation models against our benchmark using standardized metrics. PANDA-PLUS-Bench addresses a critical gap in foundation model evaluation by providing a purpose-built resource for robustness assessment in the clinically important context of Gleason grading.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。