构建首个跨模态医学影像分割测试时自适应基准,统一评估20种方法性能。
A Large Scale Benchmark for Test Time Adaptation Methods in Medical Image Segmentation
- 统一数据预处理与测试协议,系统对比4类自适应范式。
- 输入级方法在轻微外观变化下最稳定,特征级与输出级提升边界分割精度。
- 发现多类方法在大中心/设备差异下性能骤降,提示临床部署需谨慎选型。
测试时自适应(Test Time Adaptation)是缓解医学图像分割中领域偏移的有前景方法,但现有评估在模态覆盖、任务多样性和方法一致性方面仍显不足。本文提出MedSeg-TTA,一个综合性基准,涵盖七种成像模态(包括MRI、CT、超声、病理、皮肤镜、OCT和胸部X光)下20种代表性自适应方法的评估,所有实验均采用统一的数据预处理、主干网络配置和测试时协议。该基准涵盖四类关键适应范式:输入级变换、特征级对齐、输出级正则化与先验估计,首次实现跨模态的系统性比较。结果表明,无单一范式在所有条件下表现最优;输入级方法在轻度外观偏移下更稳定,特征级与输出级方法在边界相关指标上优势显著,而基于先验的方法表现出强模态依赖性。部分方法在大规模中心间及设备间偏移下性能显著下降,凸显了临床部署中方法选择的严谨性。MedSeg-TTA提供标准化数据集、验证过的代码实现及公开排行榜,为未来鲁棒且临床可靠的测试时自适应研究奠定坚实基础。所有源代码与开源数据集可于https://github.com/wenjing-gg/MedSeg-TTA获取。
原文摘要 · Abstract (English)
Test time Adaptation is a promising approach for mitigating domain shift in medical image segmentation; however, current evaluations remain limited in terms of modality coverage, task diversity, and methodological consistency. We present MedSeg-TTA, a comprehensive benchmark that examines twenty representative adaptation methods across seven imaging modalities, including MRI, CT, ultrasound, pathology, dermoscopy, OCT, and chest X-ray, under fully unified data preprocessing, backbone configuration, and test time protocols. The benchmark encompasses four significant adaptation paradigms: Input-level Transformation, Feature-level Alignment, Output-level Regularization, and Prior Estimation, enabling the first systematic cross-modality comparison of their reliability and applicability. The results show that no single paradigm performs best in all conditions. Input-level methods are more stable under mild appearance shifts. Feature-level and Output-level methods offer greater advantages in boundary-related metrics, whereas prior-based methods exhibit strong modality dependence. Several methods degrade significantly under large inter-center and inter-device shifts, which highlights the importance of principled method selection for clinical deployment. MedSeg-TTA provides standardized datasets, validated implementations, and a public leaderboard, establishing a rigorous foundation for future research on robust, clinically reliable test-time adaptation. All source codes and open-source datasets are available at https://github.com/wenjing-gg/MedSeg-TTA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。