用细胞层叠结构让病理图像注意力更准,同时生成可解释的定位图。
Aligning Cellular Sheaves with Classifier Attention for Interpretable Weakly-Supervised Pathology Localization

- 引入细胞层叠模型,通过图结构中的局部一致性检测来优化注意力。
- 在Camelyon16上实现0.940的补丁级AUC,注意力评分提升至0.953。
- 结果可跨数据集迁移,适合需要可解释性病理诊断的临床场景。
基于基础特征的注意力多实例学习(ABMIL)在整张切片分类上已接近饱和,但其注意力图存在定位不准问题。本文提出使用细胞层叠结构,在图的顶点和边处配置有限维向量空间及一致线性映射,以检测图结构数据中的局部不一致性。将层叠不一致场与ABMIL结合,设计注意条件一致性损失,利用分类器注意力定义邻近区域应一致的范围。联合训练下,不一致场在Camelyon16上达到0.940的补丁级AUC,注意力得分从0.717提升至0.953。冻结分类器的两阶段消融实验仅得0.727的不一致场得分,说明性能提升源于双目标协同优化。模型无需重训即可迁移到Camelyon17,保持0.932±0.083的ΔAUC和0.955±0.099的注意力AUC。最终输出的注意力图与层叠不一致图同步激活于诊断区域,为每张切片提供双重解释。
原文摘要 · Abstract (English)
Weakly-supervised classification of whole-slide images with attention-based multiple instance learning (ABMIL) on top of foundation features now reaches near-saturation on Camelyon16 slide-level performance, but the corresponding attention maps are an imperfect localization signal: in clinical interpretation, a model that classifies correctly without firing on the actual lesion is hard to trust. We address this gap with cellular sheaves, which equip each vertex and edge of a graph with a finite-dimensional vector space and consistent linear maps between them, providing a principled way to detect local disagreement on graph-structured data. We apply cellular sheaves to weakly-supervised tumour localization on whole-slide images, combining a sheaf disagreement field with ABMIL. The natural training objective, encouraging consistency between similar features, produces a disagreement field that tracks tissue-level texture rather than diagnostic content. We propose attention-conditional consistency, which uses the classifier's attention to define which neighbouring patches should agree. Joint training of the classifier and the sheaf under this objective produces a disagreement field with patch-level AUC 0.940 on Camelyon16 and raises the attention head from its ABMIL-alone level of 0.717 to 0.953. Two-stage ablation with the classifier frozen at its ABMIL values reaches only 0.727 on the disagreement field and leaves attention at 0.717, confirming that the gain comes from the projector co-adapting under both objectives, not from the loss change in isolation. The trained model transfers without retraining to annotated slides from Camelyon17, maintaining Delta AUC 0.932 +/- 0.083 and attention AUC 0.955 +/- 0.099. The result is an attention map and a sheaf-disagreement map that fire on the same diagnostic regions, giving clinicians two complementary explanations for each slide-level prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。