将大模型压缩成小模型,保持高精度同时大幅降低计算开销。
Distilling foundation models for robust and efficient models in digital pathology
- 用知识蒸馏技术将大模型压缩为小型模型,参数量降几个数量级。
- 在HEST和EVA基准上分别获第3和第5名,性能接近大模型。
- 对染色和扫描差异有强鲁棒性,适合实际病理诊断场景。
近年来,数字病理学中的基础模型(FM)依赖于大规模预训练数据集和模型规模的扩展,虽显著提升下游任务性能,但也带来高昂的计算成本和推理延迟。本文探索将大型基础模型蒸馏为更小模型的方法,使参数量减少数个数量级。借助蒸馏技术,我们提出的轻量化模型H0-mini在推理成本显著降低的同时,实现了与大型基础模型近乎相当的性能。该模型在多个公开基准上进行了评估,在HEST基准上取得第3名,在EVA基准上位列第5。此外,基于PLISM数据集的鲁棒性分析表明,该模型对染色和扫描条件变化具有优异适应性,显著优于现有先进模型。这为设计高效且鲁棒的轻量级数字病理模型开辟了新路径,无需牺牲性能。
原文摘要 · Abstract (English)
In recent years, the advent of foundation models (FM) for digital pathology has relied heavily on scaling the pre-training datasets and the model size, yielding large and powerful models. While it resulted in improving the performance on diverse downstream tasks, it also introduced increased computational cost and inference time. In this work, we explore the distillation of a large foundation model into a smaller one, reducing the number of parameters by several orders of magnitude. Leveraging distillation techniques, our distilled model, H0-mini, achieves nearly comparable performance to large FMs at a significantly reduced inference cost. It is evaluated on several public benchmarks, achieving 3rd place on the HEST benchmark and 5th place on the EVA benchmark. Additionally, a robustness analysis conducted on the PLISM dataset demonstrates that our distilled model reaches excellent robustness to variations in staining and scanning conditions, significantly outperforming other state-of-the art models. This opens new perspectives to design lightweight and robust models for digital pathology, without compromising on performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。