用扩散模型+解剖引导提升3D医学图像分割的泛化能力
Medical Semantic Segmentation with Diffusion Pretrain
- 通过扩散模型与解剖坐标预测联合预训练,学习通用特征
- 在13类器官分割任务中比现有恢复式预训练高7.5%的Dice分数
- 适合需要强空间感知的3D医学图像分割研究者使用
深度学习进展表明,学习鲁棒特征表示对计算机视觉任务(包括医学图像分割)的成功至关重要。尽管基于Transformer和卷积的架构均受益于预训练中的辅助任务,但3D医学图像领域对此探索较少,且面临学习通用特征表示的挑战。本文提出一种基于扩散模型并结合解剖引导的新预训练策略,针对3D医学图像数据特性设计。引入辅助扩散过程以预训练模型,生成可用于多种下游分割任务的通用特征表示。同时采用额外模型预测3D全身部位坐标,在扩散过程中提供空间引导,增强生成表示的空间感知能力,缓解定位不准问题,并提升对复杂解剖结构的理解。在13类器官分割任务上的实证验证表明,该方法相比现有恢复式预训练提升7.5%的性能,在非线性评估场景中达到67.8的平均Dice系数,与当前最优对比式预训练方法相当。
原文摘要 · Abstract (English)
Recent advances in deep learning have shown that learning robust feature representations is critical for the success of many computer vision tasks, including medical image segmentation. In particular, both transformer and convolutional-based architectures have benefit from leveraging pretext tasks for pretraining. However, the adoption of pretext tasks in 3D medical imaging has been less explored and remains a challenge, especially in the context of learning generalizable feature representations. We propose a novel pretraining strategy using diffusion models with anatomical guidance, tailored to the intricacies of 3D medical image data. We introduce an auxiliary diffusion process to pretrain a model that produce generalizable feature representations, useful for a variety of downstream segmentation tasks. We employ an additional model that predicts 3D universal body-part coordinates, providing guidance during the diffusion process and improving spatial awareness in generated representations. This approach not only aids in resolving localization inaccuracies but also enriches the model's ability to understand complex anatomical structures. Empirical validation on a 13-class organ segmentation task demonstrate the effectiveness of our pretraining technique. It surpasses existing restorative pretraining methods in 3D medical image segmentation by $7.5\%$, and is competitive with the state-of-the-art contrastive pretraining approach, achieving an average Dice coefficient of 67.8 in a non-linear evaluation scenario.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。