arXiv:2608.15970cs.CV2026-08

研究滑片选择策略如何改变模型看到的证据,揭示部署时的潜在偏差。

BagShift: Measuring How Patch Selection Changes the Evidence Seen by Whole-Slide MIL

论文配图:BagShift: Measuring How Patch Selection Changes the Evidence Seen by Whole-Slide MIL
图 1 · 摘自论文原文
  • 设计配对实验,固定特征和预测器,仅改变补丁选择方式。
  • 相同128个补丁下,局部聚焦使PANDA的QWK下降17.96点,显著低于随机采样。
  • 适用于关注模型鲁棒性与部署影响的医学图像分析研究者。

全滑片多实例学习(MIL)仅能观察其选择器所采纳的补丁。部署阶段可能因计算限制、组织掩码或区域工作流程改变选择器,即使补丁数量不变。我们提出BagShift,一种配对协议,在保持特征和预测器不变的前提下,改变同一病例的选择器,从而隔离选择器响应与病例分布的影响。在相同128补丁预算下,跨组织采样与集中于某坐标区域暴露明显不同的证据:在PANDA数据集上,两种视角分别导致四次加权κ(QWK)下降1.57和17.96点(以×100计)。在CAMELYON16上,训练中未包含的病变标注显示,局部视图仅在10.0%的微转移观察中保留肿瘤,且匹配暴露无法一致恢复损失。相同固定数量的压力源在外部肺亚型任务上产生较小响应,但相对覆盖差异使跨任务严重性描述具有意义。当可重复局部观测可用时,先合并其补丁再进行一次非线性MIL处理,使PANDA的QWK提升7.87点,优于平均区域预测。补丁数量定义的是计算量,而非实际观察证据;部署评估应同时报告选择器保留的内容及重复观测的聚合方式。

原文摘要 · Abstract (English)

Whole-slide multiple-instance learning (MIL) observes only the patches admitted by its selector. Deployment can alter this selector through compute limits, tissue masking, or regional workflows, even when the patch count is unchanged. We introduce BagShift, a paired protocol that changes the selector for the same case while holding its features and predictor fixed, thereby isolating selector response from case mix. With equal 128-patch budgets, sampling across the tissue or concentrating around one coordinate exposes markedly different evidence: on PANDA, the two views reduce quadratic weighted kappa by 1.57 and 17.96 points, respectively (QWK reported on the $\times100$ scale). On CAMELYON16, lesion annotations withheld from model development show that localized views retain tumor in only 10.0\% of micrometastatic observations, and matched exposure does not consistently recover the loss. The same fixed-count stressor produces a much smaller response on external lung subtyping, although differences in relative coverage make cross-task severity descriptive. When repeated localized observations are available, unioning their patches before one nonlinear MIL pass improves PANDA QWK by 7.87 points over averaging regional predictions. Patch count specifies computation, not observed evidence; deployment evaluations should report both what a selector preserves and how repeated observations are aggregated.

医学图像多实例学习滑片分析偏差评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。