让病理图像模型持续学习,避免遗忘旧知识并提升诊断准确率。
Lifelong Whole Slide Image Analysis: Online Vision-Language Adaptation and Past-to-Present Gradient Distillation
- 用视觉语言模型连接组织特征与文本原型,实现在线学习。
- 通过梯度蒸馏减少新旧任务间的知识遗忘,最高提升5.068%准确率。
- 适合需要长期更新的医疗图像分析系统,尤其适用于多机构协作。
全幻灯片图像(WSI)在癌症精准诊断与预后中至关重要,因其提供细胞级组织细节。然而,随着计算任务快速增长,其千兆像素级大小带来了存储、处理和训练的巨大挑战。因此,亟需为WSI分析开发持续学习方法。当幻灯片分布于多个机构时,我们旨在构建统一的在线模型,作为临床与医院场景中的癌症诊断工具。本文提出ADaFGrad方法,首先利用病理视觉-语言基础模型,建立区域组织特征与预定义文本原型缓冲区之间的交互框架;其次提出一种梯度蒸馏机制,在持续学习设置下模拟分类头参数对logit的梯度,以保持历史任务知识。我们在六个TCGA数据集序列上进行训练与评估。实验表明,仅经过少数训练轮次,ADaFGrad即超越现有WSI专用及通用持续学习方法,类增量学习场景下最高提升5.068%,且遗忘最少。同时,相比基线模型,准确率最高提升40.084%,验证了所提模块的有效性。
原文摘要 · Abstract (English)
Whole Slide Images (WSIs) play a crucial role in accurate cancer diagnosis and prognosis, as they provide tissue details at the cellular level. However, the rapid growth of computational tasks involving WSIs poses significant challenges. Given that WSIs are gigapixels in size, they present difficulties in terms of storage, processing, and model training. Therefore, it is essential to develop lifelong learning approaches for WSI analysis. In scenarios where slides are distributed across multiple institutes, we aim to leverage them to develop a unified online model as a computational tool for cancer diagnosis in clinical and hospital settings. In this study, we introduce ADaFGrad, a method designed to enhance lifelong learning for whole-slide image (WSI) analysis. First, we leverage pathology vision-language foundation models to develop a framework that enables interaction between a slide's regional tissue features and a predefined text-based prototype buffer. Additionally, we propose a gradient-distillation mechanism that mimics the gradient of a logit with respect to the classification-head parameters across past and current iterations in a continual-learning setting. We construct a sequence of six TCGA datasets for training and evaluation. Experimental results show that ADaFGrad outperforms both state-of-the-art WSI-specific and conventional continual-learning methods after only a few training epochs, exceeding them by up to +5.068% in the class-incremental learning scenario while exhibiting the least forgetting (i.e., retaining the most knowledge from previous tasks). Moreover, ADaFGrad surpasses its baseline by as much as +40.084% in accuracy, further demonstrating the effectiveness of the proposed modules.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。