ALICE将多种病理模型能力融合,打造通用病理基础模型。
ALICE: Learning a General-Purpose Pathology Foundation Model from Vision, Vision-Language, and Slide-Level Experts

- 通过多阶段聚合蒸馏,整合八类专业模型知识
- 在21个任务场景中表现优于同类模型,平均排名最高
- 适合需要跨尺度、多模态病理分析的研究与临床应用
基础模型正在重塑计算病理学,但其能力受限于预训练目标、数据来源和空间尺度,导致互补专长分散在不同主干网络中。本文提出ALICE,一种通过多阶段聚合蒸馏训练的统一基础模型,将八个仅视觉、视觉语言及全切片级教师模型的知识逐步蒸馏到单一主干的专用模块中。ALICE在24,985,184张病灶级别图像和155,604张高分辨率图像上预训练,并在21个任务场景、96项下游任务和48个数据源上评估,涵盖组织区域分析、视觉语言多模态评估和全切片临床评估。在所有三类评估设置中,ALICE在匹配任务的基础模型中均取得最佳平均排名。结果表明,聚合蒸馏可将专业化模型的互补能力整合至统一主干,实现广泛计算病理应用。模型开源地址:https://github.com/WonderLandxD/ALICE。
原文摘要 · Abstract (English)
Foundation models are reshaping computational pathology, yet their capabilities remain shaped by pretraining objectives, data sources, and spatial scales, fragmenting complementary expertise across separate backbones. Here we present ALICE, a unified foundation model trained through multi-stage agglomerative distillation that sequentially distills eight vision-only, vision-language, and slide-level teacher models into dedicated modules of a single backbone. ALICE is pretrained on 24,985,184 tile-level pathology images and 155,604 high-resolution images, and evaluated across 21 task scenarios, 96 downstream tasks, and 48 data sources, spanning region-of-interest tissue analysis, vision-language multimodal evaluation, and whole-slide clinical assessment. In all three evaluation settings, ALICE achieved the best average rank among task-matched pathology foundation models. These results demonstrate that agglomerative distillation can consolidate complementary capabilities from specialized models into a unified backbone for broad computational pathology applications. The model is available at https://github.com/WonderLandxD/ALICE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。