对比零样本分类与持续学习,探索病理图像长期学习新方案
ZeroSlide: Is Zero-Shot Classification Adequate for Lifelong Learning in Whole-Slide Image Analysis in the Era of Pathology Vision-Language Foundation Models?
- 用零样本视觉语言模型直接分类,避免重复训练
- 在多个病理任务上验证零样本方法效果,发现性能仍可提升
- 为临床部署提供轻量级长期学习思路,适合资源受限场景
全切片图像(WSI)的长期学习面临挑战:需训练统一模型以持续完成癌症分型、肿瘤分类等多任务,且要求分布式、增量式学习。由于WSI数据量大、存储和处理耗时,每次新增任务都重新训练模型效率低下。现有研究采用正则化与回放策略应对此问题。然而,随着视觉-语言基础模型能对齐病理图像与诊断文本,一个关键问题浮现:仅依赖零样本分类是否足以支持长期学习?还是需要进一步研究持续学习策略以提升性能?据我们所知,这是首个系统比较传统持续学习方法与视觉-语言零样本分类在WSI上的研究。源代码与实验结果将很快公开。
原文摘要 · Abstract (English)
Lifelong learning for whole slide images (WSIs) poses the challenge of training a unified model to perform multiple WSI-related tasks, such as cancer subtyping and tumor classification, in a distributed, continual fashion. This is a practical and applicable problem in clinics and hospitals, as WSIs are large, require storage, processing, and transfer time. Training new models whenever new tasks are defined is time-consuming. Recent work has applied regularization- and rehearsal-based methods to this setting. However, the rise of vision-language foundation models that align diagnostic text with pathology images raises the question: are these models alone sufficient for lifelong WSI learning using zero-shot classification, or is further investigation into continual learning strategies needed to improve performance? To our knowledge, this is the first study to compare conventional continual-learning approaches with vision-language zero-shot classification for WSIs. Our source code and experimental results will be available soon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。