无需训练即可融合通用与专用模型,提升医学图像分割泛化能力
MedSAMix: A Training-Free Model Merging Approach for Medical Image Segmentation
- 通过零阶优化自动寻找最佳层间融合方案
- 在25个任务上实现6.67%专项性能提升和4.37%跨任务泛化提升
- 支持单任务与多目标优化,适配不同临床场景需求
通用医学图像分割模型因其在多样任务间的强泛化能力而成为有前景的范式,展现出广泛临床应用潜力。这一潜力部分源于通用视觉模型如分割一切模型(SAM)的成功,激发了多种针对医学分割任务的微调变体发展。然而,像MedSAM这样的微调变体通常仅在有限且异构、标注稀疏、分布偏移严重的医学影像数据上训练,限制了其在广泛医学分割任务中的泛化能力。为此,我们提出MedSAMix,一种无需训练的模型融合方法,将通用模型(如SAM)与专用模型(如MedSAM)的优势结合。不同于依赖人工配置、常导致次优结果的传统融合方法,我们采用零阶优化方法自动发现最优层级融合方案。此外,为满足临床应用中对领域特异性与泛化性的不同需求,我们设计了两种策略:单任务优化与多目标优化。在25个医学分割任务上的大量评估表明,MedSAMix有效缓解了模型偏差,在领域特定准确率和泛化性能上均持续提升,专项任务提升达6.67%,多任务评估提升4.37%。
原文摘要 · Abstract (English)
Universal medical image segmentation models have emerged as a promising paradigm due to their strong generalizability across diverse tasks, showing great potential for a wide range of clinical applications. This potential has been partly driven by the success of general-purpose vision models such as the Segment Anything Model (SAM), which has inspired the development of various fine-tuned variants for medical segmentation tasks. However, fine-tuned variants like MedSAM are trained on comparatively limited medical imaging data that often suffers from heterogeneity, scarce annotations, and distributional shifts. These challenges limit their ability to generalize across a wide range of medical segmentation tasks. In this regard, we propose MedSAMix, a training-free model merging method that integrates the strengths of both generalist models (e.g., SAM) and specialist models (e.g., MedSAM) for medical image segmentation. In contrast to traditional model merging approaches that rely on manual configuration and often result in suboptimal outcomes, we propose a zero-order optimization method to automatically discover optimal layer-wise merging solutions. Furthermore, for clinical applications, we develop two regimes to meet the demand of domain-specificity and generalizability in different scenarios by single-task optimization and multi-objective optimization respectively. Extensive evaluations on 25 medical segmentation tasks demonstrate that MedSAMix effectively mitigates model bias and consistently improves performance in both domain-specific accuracy and generalization, achieving improvements of 6.67% on specialized tasks and 4.37% on multi-task evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。