100个U-Net变体在28个数据集上全面评测,找出最优选择
U-Bench: A Comprehensive Understanding of U-Net through 100-Variant Benchmarking
- 构建100种变体在28个数据集上系统评估,兼顾性能与效率
- 提出新指标U-Score,统一衡量模型表现与计算开销
- 提供智能推荐工具和开源资源,助力科研高效选型
过去十年中,U-Net是医学图像分割的主导架构,催生了数千种U形变体。然而,由于缺乏充分的统计验证及对效率与跨数据集泛化能力的考量,至今尚无全面基准来系统评估其性能与实用性。为此,我们提出了U-Bench——首个大规模、统计严谨的基准,对100种U-Net变体在28个数据集和10种成像模态上进行评估。贡献包括:(1) 全面评估:从统计鲁棒性、零样本泛化和计算效率三个维度出发,引入新型指标U-Score,联合捕捉性能与效率权衡;(2) 系统分析与选型指导:基于大规模评估结果,系统分析数据集特性与架构范式对性能的影响,并提出模型顾问代理,辅助研究者为特定任务选择最优模型;(3) 开放共享:公开全部代码、模型、协议与权重,支持复现与扩展。U-Bench不仅揭示了以往评估的不足,更为未来十年基于U-Net的分割模型建立了公平、可复现且实用的基准体系。
原文摘要 · Abstract (English)
Over the past decade, U-Net has been the dominant architecture in medical image segmentation, leading to the development of thousands of U-shaped variants. Despite its widespread adoption, there is still no comprehensive benchmark to systematically evaluate their performance and utility, largely because of insufficient statistical validation and limited consideration of efficiency and generalization across diverse datasets. To bridge this gap, we present U-Bench, the first large-scale, statistically rigorous benchmark that evaluates 100 U-Net variants across 28 datasets and 10 imaging modalities. Our contributions are threefold: (1) Comprehensive Evaluation: U-Bench evaluates models along three key dimensions: statistical robustness, zero-shot generalization, and computational efficiency. We introduce a novel metric, U-Score, which jointly captures the performance-efficiency trade-off, offering a deployment-oriented perspective on model progress. (2) Systematic Analysis and Model Selection Guidance: We summarize key findings from the large-scale evaluation and systematically analyze the impact of dataset characteristics and architectural paradigms on model performance. Based on these insights, we propose a model advisor agent to guide researchers in selecting the most suitable models for specific datasets and tasks. (3) Public Availability: We provide all code, models, protocols, and weights, enabling the community to reproduce our results and extend the benchmark with future methods. In summary, U-Bench not only exposes gaps in previous evaluations but also establishes a foundation for fair, reproducible, and practically relevant benchmarking in the next decade of U-Net-based segmentation models. The project can be accessed at: https://fenghetan9.github.io/ubench. Code is available at: https://github.com/FengheTan9/U-Bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。