arXiv:2601.16060cs.CV2026-01中稿 · IEEE ISBI被引 1

用自然语言控制扩散模型,实现多器官医学图像分割

ProGiDiff: Prompt-Guided Diffusion-Based Medical Image Segmentation

  • 通过自定义编码器控制预训练扩散模型生成分割图
  • 在CT数据上表现优于现有方法,支持多类分割
  • 可低秩微调迁移至MRI图像,适合专家协作交互

现有医学图像分割方法虽高效,但多为确定性模型,难以接受自然语言提示,缺乏多方案生成、人机交互和跨模态适应能力。最近的文本到图像扩散模型虽有潜力,但从头训练需大量数据,且常仅支持二分类,无法受语言提示控制。为此,我们提出新框架ProGiDiff,利用预训练图像生成模型进行医学图像分割。具体采用类似ControlNet的条件机制,结合自定义编码器,实现对图像的条件控制,引导扩散模型输出分割掩码。该方法可简单通过提示目标器官扩展至多类别分割。在CT图像器官分割实验中,性能优于以往方法,且在专家参与下能生成多个备选方案。更重要的是,所学条件机制可通过低秩、少样本微调,轻松迁移到MR图像分割任务。

原文摘要 · Abstract (English)

Widely adopted medical image segmentation methods, although efficient, are primarily deterministic and remain poorly amenable to natural language prompts. Thus, they lack the capability to estimate multiple proposals, human interaction, and cross-modality adaptation. Recently, text-to-image diffusion models have shown potential to bridge the gap. However, training them from scratch requires a large dataset-a limitation for medical image segmentation. Furthermore, they are often limited to binary segmentation and cannot be conditioned on a natural language prompt. To this end, we propose a novel framework called ProGiDiff that leverages existing image generation models for medical image segmentation purposes. Specifically, we propose a ControlNet-style conditioning mechanism with a custom encoder, suitable for image conditioning, to steer a pre-trained diffusion model to output segmentation masks. It naturally extends to a multi-class setting simply by prompting the target organ. Our experiment on organ segmentation from CT images demonstrates strong performance compared to previous methods and could greatly benefit from an expert-in-the-loop setting to leverage multiple proposals. Importantly, we demonstrate that the learned conditioning mechanism can be easily transferred through low-rank, few-shot adaptation to segment MR images.

医学图像扩散模型分割自然语言控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。