arXiv:2512.17774eess.IVcs.AI2025-12被引 3

改进3D ConvNeXt架构,实现医学图像分割的超大规模预训练新标杆。

MedNeXt-v2: Scaling 3D ConvNeXts for Large-Scale Supervised Representation Learning in Medical Image Segmentation

  • 设计新型3D ConvNeXt-v2,融合全局响应归一化与多尺度扩展。
  • 在1.8万例CT数据上预训练,6个基准任务平均提升3.2%以上。
  • 适合追求高精度医学影像分割的科研与临床团队使用。

大规模监督预训练正快速重塑3D医学图像分割领域。然而,现有工作主要关注数据规模扩大,忽视了骨干网络在大规模下是否具备有效表征学习能力。本文重新审视基于ConvNeXt的立体分割架构,提出MedNeXt-v2——一种复合缩放的3D ConvNeXt,通过改进微架构与数据扩展实现顶尖性能。首先,我们证明现有大规模预训练中常用骨干网络常非最优;随后,通过全面的骨干网络基准测试发现,从零开始的更强性能可可靠预测预训练后的下游表现。基于此,引入3D全局响应归一化模块,并采用深度、宽度与上下文缩放策略优化架构。我们在18,000个CT体积数据上预训练MedNeXt-v2,微调后在六个挑战性CT和MR基准(共144个结构)上达到当前最佳表现,显著优于七个公开发布的预训练模型。此外,基准分析表明:更强骨干带来更好结果,表示学习缩放对病理分割收益更大,而模态特定预训练在完整微调后几乎无益。结论:MedNeXt-v2成为3D医学图像分割中大规模监督表征学习的强力骨干。代码与预训练模型已开源至nnUNet官方仓库。

原文摘要 · Abstract (English)

Large-scale supervised pretraining is rapidly reshaping 3D medical image segmentation. However, existing efforts focus primarily on increasing dataset size and overlook the question of whether the backbone network is an effective representation learner at scale. In this work, we address this gap by revisiting ConvNeXt-based architectures for volumetric segmentation and introducing MedNeXt-v2, a compound-scaled 3D ConvNeXt that leverages improved micro-architecture and data scaling to deliver state-of-the-art performance. First, we show that routinely used backbones in large-scale pretraining pipelines are often suboptimal. Subsequently, we use comprehensive backbone benchmarking prior to scaling and demonstrate that stronger from scratch performance reliably predicts stronger downstream performance after pretraining. Guided by these findings, we incorporate a 3D Global Response Normalization module and use depth, width, and context scaling to improve our architecture for effective representation learning. We pretrain MedNeXt-v2 on 18k CT volumes and demonstrate state-of-the-art performance when fine-tuning across six challenging CT and MR benchmarks (144 structures), showing consistent gains over seven publicly released pretrained models. Beyond improvements, our benchmarking of these models also reveals that stronger backbones yield better results on similar data, representation scaling disproportionately benefits pathological segmentation, and that modality-specific pretraining offers negligible benefit once full finetuning is applied. In conclusion, our results establish MedNeXt-v2 as a strong backbone for large-scale supervised representation learning in 3D Medical Image Segmentation. Our code and pretrained models are made available with the official nnUNet repository at: https://www.github.com/MIC-DKFZ/nnUNet

3D分割预训练医学影像ConvNeXt

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。