提出新方法在不重训练模型前提下,找出能解释病理图像分类结果的少量关键区域。
Are Compact Rationales Free? Measuring Tile Selection Headroom in Frozen WSI-MIL

- 设计轻量级读出层FOCI,从冻结的分类器中提取紧凑且输出一致的候选区域。
- 在三个数据集上验证,相比原始注意力机制,所需关键区域减少32%-56%。
- 适用于需要可解释性审计的医疗图像分析场景,尤其适合模型冻结后分析。
全切片图像(WSI)多实例学习(MIL)分类器虽能达到高滑块级AUC,但整体预测过程仍不透明。注意力分数常被用作事后解释,但高注意力可能反映聚合偏好而非模型充分的紧凑理由。本文研究冻结的WSI-MIL模型的事后理由突出:给定已训练分类器,能否在不重训练主干网络的前提下,仅通过一个紧凑、输出一致的切片子集恢复其滑块级预测?我们以寻找最优上下文实例(FOCI)为例,构建一个轻量级的推理读出层,对冻结的MIL主干进行训练,目标是保证模型输出充分性与排除性,评估采用适配于WSI-MIL的插入式顺序揭示协议(SRP),并以选择余量指数(SHI)总结。在三个WSI基准和七个MIL主干上,结果显示紧凑理由的存在依赖于选择余量:变压器和多分支注意力聚合器可容纳紧凑理由,而接近最小注意力池化基线进入选择饱和状态,硬选择主干则可能与外部读出冲突。对于TransMIL,相比其记录的CLS代理排名,FOCI在各基准上将最小充分K(MSK)切片数降低32%-56%,而ACMIL+FOCI达到最高均值SHI(+0.465)。基于删除的扰动和仅使用选中切片的下游评估提供互补验证。这些结果表明,FOCI可作为模型级可解释性与审计层:所选切片并非临床或病理科级别的诊断充分性声明,而是可审查的候选理由,展示了冻结的MIL预测何时可定位到小规模输出一致子集。
原文摘要 · Abstract (English)
Whole-slide image (WSI) multiple instance learning (MIL) classifiers can achieve strong slide-level AUC while leaving the full-bag prediction opaque. Attention scores are widely reused as post-hoc explanations, but high attention can reflect aggregation preference rather than a compact, model-sufficient rationale. We study post-hoc rationale highlighting for frozen WSI-MIL: given a trained classifier, can its slide-level prediction be recovered from a compact, output-consistent tile subset without retraining the backbone? We instantiate this with Finding Optimal Contextual Instances (FOCI), a lightweight rationale-readout layer over a frozen MIL backbone. FOCI is trained with model-output sufficiency and exclusion objectives over keep/drop tile subsets, evaluated with an insertion-style Sequential Reveal Protocol (SRP) adapted to WSI-MIL, and summarized by the Selection Headroom Index (SHI). Across three WSI benchmarks and seven MIL backbones, FOCI reveals that compact rationales are selection-headroom dependent: transformer and multi-branch attention aggregators can admit compact rationales, near-minimal attention-pooling baselines enter a selection-saturation regime, and hard-selection backbones can conflict with an external readout. For TransMIL, relative to its documented CLS-proxy ranking, FOCI reduces the Minimum Sufficient K (MSK) tile count by 32-56% across benchmarks, while ACMIL+FOCI attains the highest mean SHI (+0.465). Deletion-based perturbation and selected-only downstream evaluation provide complementary checks. These results position FOCI as a model-level interpretability and audit layer: selected tiles are not claims of clinical or pathologist-level diagnostic sufficiency, but candidate rationales that offer a compact, reviewable view of when a frozen MIL prediction can be localized to a small output-consistent subset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。