用多教师蒸馏压缩病理大模型,体积减90%仍保持高精度。
Pathryoshka: Compressing Pathology Foundation Models via Multi-Teacher Knowledge Distillation with Nested Embeddings
- 多教师知识蒸馏结合嵌套嵌入,实现可调维度的模型压缩。
- 模型规模缩小86%-92%,在10个病理数据集上表现持平甚至更优。
- 适合资源有限的医院或实验室部署,推动病理AI普及。
病理基础模型(FMs)在计算病理学中取得了显著进展,但这些高性能模型常超过十亿参数,生成高维嵌入,限制了在计算资源受限场景下的研究与临床应用。本文提出Pathryoshka,一种受RADIO蒸馏和马特里什卡表征学习启发的多教师知识蒸馏框架,可在保持可变嵌入维度的同时压缩病理模型。我们在包含十个不同下游任务的公开病理基准上评估该框架,结果表明,相比其大型教师模型,Pathryoshka将模型规模减少86%-92%,性能相当;且在同等规模下,相比最先进的单教师蒸馏模型,准确率中位数提升7.0。该方法使高性能病理模型能在本地高效部署,兼顾准确性与表征丰富性,助力更广泛的科研与临床群体使用前沿病理模型。
原文摘要 · Abstract (English)
Pathology foundation models (FMs) have driven significant progress in computational pathology. However, these high-performing models can easily exceed a billion parameters and produce high-dimensional embeddings, thus limiting their applicability for research or clinical use when computing resources are tight. Here, we introduce Pathryoshka, a multi-teacher distillation framework inspired by RADIO distillation and Matryoshka Representation Learning to reduce pathology FM sizes while allowing for adaptable embedding dimensions. We evaluate our framework with a distilled model on ten public pathology benchmarks with varying downstream tasks. Compared to its much larger teachers, Pathryoshka reduces the model size by 86-92% at on-par performance. It outperforms state-of-the-art single-teacher distillation models of comparable size by a median margin of 7.0 in accuracy. By enabling efficient local deployment without sacrificing accuracy or representational richness, Pathryoshka democratizes access to state-of-the-art pathology FMs for the broader research and clinical community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。