病理图像分析中,模型常忽视空间位置,导致误判。
Spatial Blindness in Whole-Slide Multiple Instance Learning

- 用原型直方图+坐标扰动约束,增强空间感知能力
- 在9个公开数据集上提升分类与生存预测性能
- 适合需要精确定位的医学图像诊断任务
全切片多实例学习(WSI MIL)模型常被称作具有上下文感知能力,但实际在组织结构关键的病理任务中,多个强基准模型在打乱切片块坐标后,整体分类AUC几乎不变。其预测准确但依赖局部外观组合,而非空间关系。这种现象称为空间盲视。我们提出基于优化机制的解释:滑片级监督下,密集外观统计早期被学习,导致稀疏空间关系梯度过弱。ResTopoMIL通过先拟合一个置换不变的原型直方图并冻结,再让轻量图分支在坐标打乱约束下学习残差来解决。该架构设计简单,关键在于空间分支的训练方式。在9个公开的全切片图像基准上,仅用115万参数,显著提升分类与生存预测性能,恢复对坐标准确性的敏感性,并在CAMELYON-16上提供更强定位证据。
原文摘要 · Abstract (English)
Whole-slide MIL models are often called context-aware once graphs, Transform ers, or state-space modules are placed above patch embeddings. We show that this label can be deceptive. On pathology tasks where tissue architecture is part of the diagnostic signal, several strong MIL baselines retain nearly unchanged slide level AUC after patch coordinates are permuted. Their predictions are accurate, but largely compositional. We refer to this failure mode as spatial blindness. Our explanation is optimization-based: dense appearance statistics are learned early under slide-level supervision, leaving weak gradients for sparse spatial relations. ResTopoMIL addresses the issue by first fitting a permutation-invariant prototype histogram and then freezing it while a lightweight graph branch learns the residual under a coordinate-shuffling constraint. The architecture is simple by design; the intervention is in how the spatial branch is trained. Across 9 public WSI bench marks, ResTopoMIL improves classification and survival prediction with 1.15M parameters, restores sensitivity to coordinate perturbation, and gives stronger lo calization evidence on CAMELYON-16.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。