arXiv:2603.12886cs.CV2026-03被引 2

提出评估病理模型抗染色差异的三步协议,助力可靠部署。

A protocol for evaluating robustness to H&E staining variation in computational pathology models

  • 构建参考染色库,分三步评估模型对染色变化的鲁棒性
  • 306个模型在不同染色条件下AUC波动达0.142,鲁棒性差距0.072
  • 发现性能越好鲁棒性越差,助选适配实际场景的模型

计算病理(CPath)模型在真实场景中受限于苏木精-伊红(H&E)染色的实验室间差异。本文提出一种三步评估协议:第一步选取参考染色条件,第二步表征测试集染色属性,第三步在模拟参考染色条件下应用模型。基于PLISM数据集构建新参考染色库,以未见的SurGen结直肠癌数据集(n=738)为例,评估306个微卫星不稳定性(MSI)分类模型,包括300个基于注意力的多实例学习模型(使用UNI2-h、H-Optimus-1、Virchow2三种特征提取器)及6个公开模型。分类性能以AUC衡量,鲁棒性定义为四种模拟染色条件(高低H&E强度、高低颜色相似度)下的最小-最大AUC差值。模型表现范围为AUC 0.769–0.911(Δ=0.142),鲁棒性范围为0.007–0.079(Δ=0.072),且与性能呈弱负相关(Pearson r=-0.22,95% CI [-0.34, -0.11])。结果表明该协议可实现鲁棒性导向的模型选择,并揭示性能在染色变化下的动态,支持模型可靠部署的范围界定。代码已开源。

原文摘要 · Abstract (English)

Sensitivity to staining variation remains a major barrier to deploying computational pathology (CPath) models as hematoxylin and eosin (H&E) staining varies across laboratories, requiring systematic assessment of how this variability affects model prediction. In this work, we developed a three-step protocol for evaluating robustness to H&E staining variation in CPath models. Step 1: Select reference staining conditions, Step 2: Characterize test set staining properties, Step 3: Apply CPath model(s) under simulated reference staining conditions. Here, we first created a new reference staining library based on the PLISM dataset. As an exemplary use case, we applied the protocol to assess the robustness properties of 306 microsatellite instability (MSI) classification models on the unseen SurGen colorectal cancer dataset (n=738), including 300 attention-based multiple instance learning models trained on the TCGA-COAD/READ datasets across three feature extractors (UNI2-h, H-Optimus-1, Virchow2), alongside six public MSI classification models. Classification performance was measured as AUC, and robustness as the min-max AUC range across four simulated staining conditions (low/high H&E intensity, low/high H&E color similarity). Across models and staining conditions, classification performance ranged from AUC 0.769-0.911 ($Δ$ = 0.142). Robustness ranged from 0.007-0.079 ($Δ$ = 0.072), and showed a weak inverse correlation with classification performance (Pearson r=-0.22, 95% CI [-0.34, -0.11]). Thus, we show that the proposed evaluation protocol enables robustness-informed CPath model selection and provides insight into performance shifts across H&E staining conditions, supporting the identification of operational ranges for reliable model deployment. Code is available at https://github.com/CTPLab/staining-robustness-evaluation .

计算病理染色鲁棒性模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。