通过证据球结构提升少样本病理图像分类的可解释性与准确率
From Patches to Evidence Balls: Class-Conditioned Evidence Retrieval for Few-Shot Whole Slide Image Classification

- 将局部切片聚合成语义空间一致的证据球,实现结构化证据组织
- 在少样本设置下超越传统方法,在4个任务上均取得更优性能
- 支持空间定位与类别专属证据,适合需要可解释性的医疗诊断场景
全切片图像(WSI)分类是依赖证据的任务,诊断线索常稀疏、空间分布且类别相关。现有基于多实例学习(MIL)和视觉-语言的方法将大量切片特征聚合为单一全局表示,在少样本监督下,有限的滑片级标签难以学习可靠的聚合机制,无法有效组织稀疏局部线索为紧凑连贯的诊断证据。此外,共享的滑片表示将支持候选类及其替代类的证据压缩在同一特征中,限制了类别特异性推理与可解释性。为此,我们提出EviBall,一种面向少样本WSI分类的类别条件证据检索框架。EviBall通过语义-空间分配与中心精炼,将局部切片组织为紧凑且空间连贯的证据球,实现弱监督下的结构化证据单元构建。随后,利用任务特定的类别查询——包括形态学导向的语言引导查询和分子终点预测的分子引导查询——检索支持性证据球,生成用于直接类别预测的类别条件证据表示。通过引入结构化证据单元与任务相关的语义引导,EviBall减少了对从稀缺滑片级标签中学习无约束全局聚合机制的依赖,从而将少样本WSI分类重新定义为结构化证据检索与候选类别间的竞争。在四个形态学导向与分子终点预测的WSI任务上的广泛实验表明,EviBall在多种少样本设置下持续优于传统MIL与视觉-语言基线,并为每次预测提供空间定位且类别专属的证据。
原文摘要 · Abstract (English)
Whole slide image (WSI) classification is an evidence-driven task, where diagnostic cues are often sparse, spatially organized, and class-dependent. Existing MIL and vision-language methods aggregate a large pool of patch features into a single global slide representation. Under few-shot supervision, limited slide-level labels make it difficult to learn a reliable aggregation mechanism that organizes sparse local cues into compact and coherent diagnostic evidence. Moreover, a shared slide representation compresses evidence supporting a candidate class and its alternatives into the same feature, limiting class-specific reasoning and interpretability. To address these issues, we propose EviBall, a class-conditioned evidence retrieval framework for few-shot WSI classification. EviBall organizes local patches into Evidence Balls through semantic-spatial assignment and center refinement, yielding compact and spatially coherent evidence units under weak supervision. It then uses task-specific class queries, including language-guided queries for morphology-oriented tasks and molecular-guided queries for molecular endpoint prediction, to retrieve supporting evidence balls and produce class-conditioned evidence representations for direct class-wise prediction. By introducing structured evidence units and task-relevant semantic guidance, EviBall reduces the reliance on learning an unconstrained global aggregation mechanism from scarce slide-level labels. It therefore reformulates few-shot WSI classification as structured evidence retrieval and competition among candidate classes. Extensive experiments across four morphology-oriented and molecular endpoint WSI tasks demonstrate that EviBall consistently outperforms conventional and vision-language MIL baselines under diverse few-shot settings, while providing spatially localized and class-specific evidence for each prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。