arXiv:2603.15967cs.CV2026-03被引 4

评估11个病理模型在肾病图像上的表现,发现其对细微结构识别能力有限。

A Comprehensive Benchmark of Histopathology Foundation Models for Kidney Digital Pathology Images

  • 系统测试11个开源模型在多种染色、尺度和任务上的表现
  • 粗粒度结构识别效果中等至良好,细粒度分析表现差
  • 适用于诊断分类,但难胜任预后判断与微结构区分

组织病理学基础模型(HFMs)在癌症病理图像上已取得进展,但在非癌性慢性肾病中的应用仍不明确。本文系统评估了11个公开可用的HFMs在11项肾脏特异性下游任务中的表现,涵盖PAS、H&E、PASM和IHC四种染色方式,以及切片级与瓦片级的多尺度分析,涉及分类、回归与拷贝数检测等多种任务类型,目标包括检测、诊断与预后。瓦片级任务采用重复分层组交叉验证,切片级任务使用重复嵌套分层交叉验证。通过弗里德曼检验结合成对威尔科克森符号秩检验及霍姆-博尼费罗尼校正进行统计显著性分析,并以紧凑字母显示可视化结果。为提升可复现性,我们发布了开源Python包kidney-hfm-eval。结果显示,模型在依赖粗粒度肾组织形态的任务中表现中等至良好,如诊断分类与显著结构异常检测;但在需要精细微结构辨识、复杂生物表型或切片级预后推断的任务中表现持续下降,且与染色方式无关。总体表明当前HFMs主要编码静态中等尺度表征,对细微肾病变化或预后信号捕捉能力有限。研究呼吁开发针对肾脏、多染色、多模态的基础模型,以支持肾病临床决策。

原文摘要 · Abstract (English)

Histopathology foundation models (HFMs), pretrained on large-scale cancer datasets, have advanced computational pathology. However, their applicability to non-cancerous chronic kidney disease remains underexplored, despite coexistence of renal pathology with malignancies such as renal cell and urothelial carcinoma. We systematically evaluate 11 publicly available HFMs across 11 kidney-specific downstream tasks spanning multiple stains (PAS, H&E, PASM, and IHC), spatial scales (tile and slide-level), task types (classification, regression, and copy detection), and clinical objectives, including detection, diagnosis, and prognosis. Tile-level performance is assessed using repeated stratified group cross-validation, while slide-level tasks are evaluated using repeated nested stratified cross-validation. Statistical significance is examined using Friedman test followed by pairwise Wilcoxon signed-rank testing with Holm-Bonferroni correction and compact letter display visualization. To promote reproducibility, we release an open-source Python package, kidney-hfm-eval, available at https://pypi.org/project/kidney-hfm-eval/ , that reproduces the evaluation pipelines. Results show moderate to strong performance on tasks driven by coarse meso-scale renal morphology, including diagnostic classification and detection of prominent structural alterations. In contrast, performance consistently declines for tasks requiring fine-grained microstructural discrimination, complex biological phenotypes, or slide-level prognostic inference, largely independent of stain type. Overall, current HFMs appear to encode predominantly static meso-scale representations and may have limited capacity to capture subtle renal pathology or prognosis-related signals. Our results highlight the need for kidney-specific, multi-stain, and multimodal foundation models to support clinically reliable decision-making in nephrology.

病理分析肾病建模基础模型医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。