提出新方法评估病理图像MIL模型热图质量,发现非注意力类解释更可靠。
Beyond Attention Heatmaps: How to Get Better Explanations for Multiple Instance Learning Models in Histopathology
- 设计无额外标签的热图评估框架,适用于多种模型与任务。
- 实验表明扰动法、LRP和积分梯度优于注意力热图,更反映决策机制。
- 验证热图可关联空间转录组并揭示HPV预测不同策略,助力生物发现。
多实例学习(MIL)推动了计算病理学发展,将海量病理切片块聚合为整体诊断预测。热图常用于验证MIL模型并发现组织生物标志物,但其有效性尚未深入研究。本文提出一种无需额外标注的通用热图评估框架,对六种解释方法在分类、回归、生存分析三类任务,以及基于注意力、Transformer、Mamba的MIL模型架构,搭配UNI2、Virchow2等编码器,在大规模基准测试中进行评估。结果表明,解释质量主要取决于模型架构与任务类型;扰动(Single)、层间相关性传播(LRP)和积分梯度(IG)始终优于注意力与梯度显著性热图,后者常无法反映真实决策机制。进一步展示最优方法的能力:(i)证明批量基因表达预测模型的热图可与空间转录组数据相关联,实现生物学验证;(ii)揭示头颈部癌切片中预测人乳头瘤病毒(HPV)感染的不同模型策略。本工作强调需验证MIL热图,并证实提升可解释性可增强模型可信度并带来生物学洞见,呼吁在数字病理中更广泛采用可解释AI。代码已开源:https://github.com/bifold-pathomics/xMIL/tree/xmil-journal
原文摘要 · Abstract (English)
Multiple instance learning (MIL) has enabled substantial progress in computational histopathology, where a large amount of patches from gigapixel whole slide images are aggregated into slide-level predictions. Heatmaps are widely used to validate MIL models and to discover tissue biomarkers. Yet, the validity of these heatmaps has barely been investigated. In this work, we introduce a general framework for evaluating the quality of MIL heatmaps without requiring additional labels. We conduct a large-scale benchmark experiment to assess six explanation methods across histopathology task types (classification, regression, survival), MIL model architectures (Attention-, Transformer-, Mamba-based), and patch encoder backbones (UNI2, Virchow2). Our results show that explanation quality mostly depends on MIL model architecture and task type, with perturbation ("Single"), layer-wise relevance propagation (LRP), and integrated gradients (IG) consistently outperforming attention-based and gradient-based saliency heatmaps, which often fail to reflect model decision mechanisms. We further demonstrate the advanced capabilities of the best-performing explanation methods: (i) We provide a proof-of-concept that MIL heatmaps of a bulk gene expression prediction model can be correlated with spatial transcriptomics for biological validation, and (ii) showcase the discovery of distinct model strategies for predicting human papillomavirus (HPV) infection from head and neck cancer slides. Our work highlights the importance of validating MIL heatmaps and establishes that improved explainability can enable more reliable model validation and yield biological insights, making a case for a broader adoption of explainable AI in digital pathology. Our code is provided in a public GitHub repository: https://github.com/bifold-pathomics/xMIL/tree/xmil-journal
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。