arXiv:2602.03887eess.IVcs.CV2026-02被引 1

17种病理基础模型在18个数据集上系统评估,揭示高效分割的选型与调优策略。

To What Extent Do Token-Level Representations from Pathology Foundation Models Improve Dense Prediction?

  • 构建统一评估框架,对比17种病理基础模型在多任务中的表现
  • 发现不同模型在异质数据上的性能差异,关键在适配策略选择
  • 提供可复现工具包,助力临床实际部署的模型选型决策

病理基础模型(PFM)快速发展,已成为下游临床任务的通用骨干,具备跨组织和组织的良好迁移能力。然而,对于密集预测任务(如分割),其实际部署仍缺乏对不同PFM在各类数据集上行为的清晰、可复现的理解,以及适配策略如何影响性能与稳定性的系统分析。本文提出PFM-DenseBench,一个大规模的密集病理预测基准,涵盖17种PFM在18个公开分割数据集上的评估。在统一协议下,系统性测试多种适配与微调策略,得出关于不同PFM及调优方式在异质数据集中成功或失败的关键实践洞察。我们公开容器、配置文件与数据卡,支持可复现评估与真实场景中对模型的明智选择。

原文摘要 · Abstract (English)

Pathology foundation models (PFMs) have rapidly advanced and are becoming a common backbone for downstream clinical tasks, offering strong transferability across tissues and institutions. However, for dense prediction (e.g., segmentation), practical deployment still lacks a clear, reproducible understanding of how different PFMs behave across datasets and how adaptation choices affect performance and stability. We present PFM-DenseBench, a large-scale benchmark for dense pathology prediction, evaluating 17 PFMs across 18 public segmentation datasets. Under a unified protocol, we systematically assess PFMs with multiple adaptation and fine-tuning strategies, and derive insightful, practice-oriented findings on when and why different PFMs and tuning choices succeed or fail across heterogeneous datasets. We release containers, configs, and dataset cards to enable reproducible evaluation and informed PFM selection for real-world dense pathology tasks. Project Website: https://m4a1tastegood.github.io/PFM-DenseBench

病理分析基础模型密集预测可复现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。