用视觉语言模型提升病理图像生存分析的可解释性与数据效率
Interpretable Vision-Language Survival Analysis with Ordinal Inductive Bias for Computational Pathology
- 引入视觉语言基础模型,以文本先验指导图像特征聚合
- 在5个数据集上实现优于现有方法的生存预测性能
- 通过归因分析实现结果可解释,适合临床辅助决策场景
组织病理全切片图像(WSI)是计算病理学中评估癌症预后的关键工具。现有生存分析方法多依赖高表达能力网络和粗粒度患者标签,难以应对训练数据稀缺及标准多实例学习(MIL)框架下的弱监督问题。本文首次提出基于视觉-语言的生存分析(VLSA)新范式:利用病理视觉-语言基础模型,不再依赖复杂网络,提升数据效率;在视觉端,引入文本预设的预后先验作为辅助信号,指导实例级特征聚合,弥补MIL弱监督缺陷。针对生存分析特性,提出有序生存提示学习,将连续生存标签转为文本提示,并采用有序发病率函数作为预测目标,使生存分析与视觉-语言预测兼容。此外,通过基于Shapley值的方法实现预测结果的直观可解释性。在五个数据集上的大量实验验证了该方案的有效性。VLSA为计算病理中的生存分析提供了一种弱监督多实例学习的新路径,能有效从千兆像素级全切片图像中挖掘有价值预后线索。源代码已公开于https://github.com/liupei101/VLSA。
原文摘要 · Abstract (English)
Histopathology Whole-Slide Images (WSIs) provide an important tool to assess cancer prognosis in computational pathology (CPATH). While existing survival analysis (SA) approaches have made exciting progress, they are generally limited to adopting highly-expressive network architectures and only coarse-grained patient-level labels to learn visual prognostic representations from gigapixel WSIs. Such learning paradigm suffers from critical performance bottlenecks, when facing present scarce training data and standard multi-instance learning (MIL) framework in CPATH. To overcome it, this paper, for the first time, proposes a new Vision-Language-based SA (VLSA) paradigm. Concretely, (1) VLSA is driven by pathology VL foundation models. It no longer relies on high-capability networks and shows the advantage of data efficiency. (2) In vision-end, VLSA encodes textual prognostic prior and then employs it as auxiliary signals to guide the aggregating of visual prognostic features at instance level, thereby compensating for the weak supervision in MIL. Moreover, given the characteristics of SA, we propose i) ordinal survival prompt learning to transform continuous survival labels into textual prompts; and ii) ordinal incidence function as prediction target to make SA compatible with VL-based prediction. Notably, VLSA's predictions can be interpreted intuitively by our Shapley values-based method. The extensive experiments on five datasets confirm the effectiveness of our scheme. Our VLSA could pave a new way for SA in CPATH by offering weakly-supervised MIL an effective means to learn valuable prognostic clues from gigapixel WSIs. Our source code is available at https://github.com/liupei101/VLSA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。