用轻量级融合方法提升病理模型诊断效果,无需重训练
Plug-and-Play Logit Fusion for Heterogeneous Pathology Foundation Models
- 将多个独立训练的病理模型输出作为专家,动态学习融合权重
- 在22个任务中20项排名第一,平均性能提升约3%
- 无需特征对齐或重新训练,训练成本降低12倍,适合快速部署
病理基础模型(FMs)已成为计算病理学的核心,可在多种诊断与预后任务中实现优异的迁移性能。然而,病理基础模型的快速涌现带来了模型选择瓶颈:单一模型并非始终最优,而为每个下游任务逐一适配和验证众多候选模型代价过高。为此,我们提出一种轻量且新颖的模型融合策略——LogitProd,将独立训练的基于FM的预测器视为固定专家,学习其全切片输出的样本自适应融合权重。该融合仅作用于原始分数(logits),无需重新训练编码器,也无需异构主干网络间的特征空间对齐。我们进一步提供理论分析,证明最优加权乘积融合在训练目标下至少等同于表现最佳的单个专家。我们在涵盖全切片分类、瓦片级分类、基因突变预测和离散时间生存建模的22个基准上系统评估了LogitProd,结果在20/22项任务中排名第一,整体平均性能较最强单个专家提升约3%。该方法使从业者能以即插即用方式升级异构病理模型流水线,以约12倍更低的训练成本获得多专家增益。
原文摘要 · Abstract (English)
Pathology foundation models (FMs) have become central to computational histopathology, offering strong transfer performance across a wide range of diagnostic and prognostic tasks. The rapid proliferation of pathology foundation models creates a model-selection bottleneck: no single model is uniformly best, yet exhaustively adapting and validating many candidates for each downstream endpoint is prohibitively expensive. We address this challenge with a lightweight and novel model fusion strategy, LogitProd, which treats independently trained FM-based predictors as fixed experts and learns sample-adaptive fusion weights over their slide-level outputs. The fusion operates purely on logits, requiring no encoder retraining and no feature-space alignment across heterogeneous backbones. We further provide a theoretical analysis showing that the optimal weighted product fusion is guaranteed to perform at least as well as the best individual expert under the training objective. We systematically evaluate LogitProd on \textbf{22} benchmarks spanning WSI-level classification, tile-level classification, gene mutation prediction, and discrete-time survival modeling. LogitProd ranks first on 20/22 tasks and improves the average performance across all tasks by ~3% over the strongest single expert. LogitProd enables practitioners to upgrade heterogeneous FM-based pipelines in a plug-and-play manner, achieving multi-expert gains with $\sim$12$\times$ lower training cost than feature-fusion alternatives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。