提出新指标衡量病理模型对非生物干扰的鲁棒性,可精准识别模型弱点。
A Distributional Robustness Margin For Pathology Foundation Models

- 设计跨混杂因子鲁棒性边际CRoMa,逐样本评估生物与非生物变异距离。
- 20个切片级编码器测试显示,中位CRoMa相关性达0.90,但所有模型均有受干扰样本。
- 高CRoMa值对应下游任务中更小的捷径损失,适合评估模型抗偏差能力。
病理基础模型会捕捉组织制备、染色和扫描引入的非生物变异,导致捷径学习,影响跨机构泛化。现有鲁棒性指数(RI)因结构缺陷难以实现跨模型可靠比较,亟需更严谨的度量标准。本文提出交叉混杂因子鲁棒性边际(CRoMa),一个有符号的、针对每一样本的度量,用于判断具有相同生物学特征但不同混杂因子的样本是否比具有相同混杂因子但不同生物学特征的样本更接近。该指标适用于每个样本,使模型可在同一队列上比较,并将鲁棒性分析为分布而非单一汇总分数。我们在三个基准上对20个切片级编码器进行了评估,中位CRoMa在不同基准间高度一致(斯皮尔曼等级相关系数~0.90),但每个编码器均存在受混杂因子主导的样本,其出现频率和严重程度差异显著。在另一个基准上对四个切片级编码器的扩展分析也呈现相似模式,表明该方法超越切片层级。中位CRoMa越高,下游线性探测器的捷径诱导性能损失越小,支持其作为表示层捷径敏感性的指示器。
原文摘要 · Abstract (English)
Pathology foundation models encode non-biological variation introduced by tissue preparation, staining and scanning, enabling shortcut learning that undermines generalisation across institutions. The Robustness Index (RI) was proposed to assess whether local representation geometry is dominated by biological or non-biological variation. However, its construction suffers from structural limitations that make cross-model comparison unreliable, calling for a more principled metric. We introduce the Cross-confounder Robustness Margin (CRoMa), a signed, per-sample margin that measures whether samples sharing the same biology but different confounder lie closer than samples sharing the same confounder but different biology. It is defined for every sample, allowing models to be compared on the same cohort and robustness to be analysed as a distribution rather than reduced to a single pooled score. We evaluated CRoMa across 20 tile-level encoders on three benchmarks. Rankings by median CRoMa were highly consistent across benchmarks (Spearman rho ~ 0.90), yet every encoder retained confounder-dominated samples, whose prevalence and severity varied markedly. Similar patterns emerged for four slide-level encoders evaluated on a separate benchmark, extending the analysis beyond tile-level representations. Higher median CRoMa was associated with smaller shortcut-induced performance losses in downstream linear probes, supporting its use as a representation-level indicator of shortcut susceptibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。