arXiv:2512.00350eess.IVcs.AI2025-12

用语义先验引导扩散模型,实现轻量高效医疗图像分割

MedCondDiff: Lightweight, Robust, Semantically Guided Diffusion for Medical Image Segmentation

  • 通过金字塔视觉变压器提取语义先验,指导去噪过程
  • 相比传统扩散模型,推理速度更快,显存占用更低
  • 在多器官多模态数据上表现稳定,适合临床部署

我们提出MedCondDiff,一种基于扩散的多器官医学图像分割框架,兼具高效性与解剖学合理性。该模型利用金字塔视觉变压器(PVT)骨干网络提取语义先验,对去噪过程进行语义引导,构建轻量级且鲁棒的扩散架构。实验表明,相较于传统扩散模型,MedCondDiff在多器官、多模态数据集上显著降低推理时间和显存占用,同时保持优异分割性能,验证了语义引导扩散模型在医学影像任务中的有效性。

原文摘要 · Abstract (English)

We introduce MedCondDiff, a diffusion-based framework for multi-organ medical image segmentation that is efficient and anatomically grounded. The model conditions the denoising process on semantic priors extracted by a Pyramid Vision Transformer (PVT) backbone, yielding a semantically guided and lightweight diffusion architecture. This design improves robustness while reducing both inference time and VRAM usage compared to conventional diffusion models. Experiments on multi-organ, multi-modality datasets demonstrate that MedCondDiff delivers competitive performance across anatomical regions and imaging modalities, underscoring the potential of semantically guided diffusion models as an effective class of architectures for medical imaging tasks.

医学图像分割扩散模型轻量化语义引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。