arXiv:2507.07860cs.CV2025-07NeurIPS被引 12

构建病理图像基准测试,评估模型在多种任务中的表现与可靠性。

THUNDER: Tile-level Histopathology image UNDERstanding benchmark

  • 设计动态基准框架,支持23种模型在16个数据集上快速对比
  • 覆盖特征空间分析、鲁棒性与不确定性评估等关键指标
  • 适合医疗AI研究者用于模型选型与性能验证

数字病理学领域近年涌现出大量基础模型,用于切片级图像的特征提取,但方法繁多且更新迅速,难以有效评估。为清晰把握研究进展,我们提出THUNDER——一个面向切片级图像的病理学基础模型基准测试框架。该框架支持在16个不同数据集上对23种主流基础模型进行高效比较,涵盖多种下游任务,深入分析特征空间,并评估模型预测的鲁棒性与不确定性。THUNDER具有快速、易用、可扩展的特点,已支持多种先进模型及用户自定义模型的直接对比。代码开源,网址:https://github.com/MICS-Lab/thunder。

原文摘要 · Abstract (English)

Progress in a research field can be hard to assess, in particular when many concurrent methods are proposed in a short period of time. This is the case in digital pathology, where many foundation models have been released recently to serve as feature extractors for tile-level images, being used in a variety of downstream tasks, both for tile- and slide-level problems. Benchmarking available methods then becomes paramount to get a clearer view of the research landscape. In particular, in critical domains such as healthcare, a benchmark should not only focus on evaluating downstream performance, but also provide insights about the main differences between methods, and importantly, further consider uncertainty and robustness to ensure a reliable usage of proposed models. For these reasons, we introduce THUNDER, a tile-level benchmark for digital pathology foundation models, allowing for efficient comparison of many models on diverse datasets with a series of downstream tasks, studying their feature spaces and assessing the robustness and uncertainty of predictions informed by their embeddings. THUNDER is a fast, easy-to-use, dynamic benchmark that can already support a large variety of state-of-the-art foundation, as well as local user-defined models for direct tile-based comparison. In this paper, we provide a comprehensive comparison of 23 foundation models on 16 different datasets covering diverse tasks, feature analysis, and robustness. The code for THUNDER is publicly available at https://github.com/MICS-Lab/thunder.

病理分析模型评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。