arXiv:2608.29677cs.CVcs.AI2026-08

医学图像分割基准测试框架,让实验条件透明可复现。

MedSegBenchmarker: A Raw-Count-First Framework for Controlled 2D Medical Image Segmentation Benchmarks

论文配图:MedSegBenchmarker: A Raw-Count-First Framework for Controlled 2D Medical Image Segmentation Benchmarks
图 1 · 摘自论文原文
  • 用配置驱动方式统一数据划分、评估协议和训练流程
  • 导出像素级预测结果,支持事后分析而不重跑模型
  • 适合需要严谨对比模型性能的研究者和审稿人

尽管医学图像分割(MIS)进展迅速,但公平且可复现的模型比较仍面临挑战,原因包括数据集异质性、评估协议不一致以及架构快速迭代。现有比较常隐含假设:模型排名不受数据划分、预处理、指标聚合、不确定性估计和计算约束影响。缺乏可扩展的统一评估框架,也限制了对新模型、数据集和训练范式系统性研究。本文提出MedSegBenchmarker(MSB),一个面向2D MIS的配置驱动控制性基准测试框架。它集成重复与近似重复图像检测、分组感知数据划分、YAML研究配置、可恢复训练、超参数优化、交叉验证及基于检查点的评估。不同于仅保留汇总性能,MSB导出样本级和类别级像素计数与预测结果,并附带评估上下文。这些基础数据支持无需重复推理的后验分析。我们在三个异构2D数据集上,使用多个MIS和通用视觉模型,在256和512像素输入分辨率下进行案例研究。相同预测结果的重新聚合导致六种设置中三种的顶级架构发生变化,尽管不同聚合策略间排名相关性较高。提升输入分辨率在不同模型和数据集上带来性能增益或损失,需结合实测推理复杂度综合考量。结果表明,评估与实验设置中的微小选择可能影响基准结论。MSB已在GitHub开源,为使基准条件与评估选择明确化、可复现提供实用且可扩展的基础。

原文摘要 · Abstract (English)

Despite rapid advances in MIS, fair and reproducible comparisons of segmentation models remain challenging due to heterogeneous datasets, inconsistent evaluation protocols, and rapidly evolving architectures. In particular, comparisons often implicitly assume that model rankings are invariant to data partitioning, preprocessing, metric aggregation, uncertainty estimation, and computational constraints. The lack of extensible and unified evaluation frameworks further limits systematic investigation of new models, datasets, and training paradigms. We present MEDSEGBENCHMARKER (MSB), a configuration-driven framework for controlled benchmarking of 2D MIS. It integrates duplicate and near-duplicate image detection, group-aware data splitting, YAML study specifications, resumable training, hyperparameter optimization, cross-validation, and checkpoint-based evaluation. Rather than retaining only aggregate performance measures, MSB exports sample- and class-level pixel counts and predictions together with the evaluation context. These elementary artifacts enable post-hoc analyses without repeated inference. We demonstrate MSB in a case study involving three heterogeneous 2D datasets and multiple MIS and general-purpose vision models evaluated at 256- and 512-pixel input resolutions. Reaggregation of identical predictions changes the top-ranked architecture in three of six dataset-resolution settings, despite high rank correlations between aggregation strategies. Increasing input resolution produces model- and dataset-dependent performance gains and losses that must be considered alongside empirically measured inference complexity. These results show that seemingly minor choices in evaluation and experimental setup can affect benchmark conclusions. MSB, available at GitHub, provides a practical and extensible basis for making benchmark conditions and evaluation choices explicit and reproducible.

医学图像分割基准可复现性框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。