arXiv:2604.03134cs.CV2026-04

用稳定扩散模型做少量标注的医学图像分割,效果更好且泛化更强。

SD-FSMIS: Adapting Stable Diffusion for Few-Shot Medical Image Segmentation

  • 利用支持集与查询图交互,结合视觉转文本条件,适配扩散模型。
  • 在标准和跨域场景下均达到领先性能,少样本下表现优异。
  • 适合医疗图像分割、小样本学习研究者参考使用。

少样本医学图像分割(FSMIS)旨在仅用少量标注样本对新类医学图像进行分割,以应对医学影像中数据稀缺与领域偏移的挑战。尽管扩散模型(DM)在视觉任务中表现卓越,但其在FSMIS中的潜力尚未充分探索。本文提出,大规模扩散模型所学习的丰富视觉先验可为更鲁棒、数据高效的分割方法提供坚实基础。我们引入SD-FSMIS框架,有效适配预训练的稳定扩散(SD)模型用于FSMIS任务。通过重构其条件生成架构,提出两个关键组件:支持-查询交互(SQI)和视觉到文本条件翻译器(VTCT)。SQI提供了一种简洁而强大的方式将SD适配至FSMIS范式;VTCT模块将支持集的视觉线索转换为隐式文本嵌入,指导扩散过程,实现精准生成条件。大量实验表明,SD-FSMIS在标准设置下达到与当前最优方法相当的结果,且在更具挑战性的跨域场景中展现出出色的泛化能力。这些发现凸显了适配大规模生成模型在推进数据高效且稳健的医学图像分割方面的巨大潜力。

原文摘要 · Abstract (English)

Few-Shot Medical Image Segmentation (FSMIS) aims to segment novel object classes in medical images using only minimal annotated examples, addressing the critical challenges of data scarcity and domain shifts prevalent in medical imaging. While Diffusion Models (DM) excel in visual tasks, their potential for FSMIS remains largely unexplored. We propose that the rich visual priors learned by large-scale DMs offer a powerful foundation for a more robust and data-efficient segmentation approach. In this paper, we introduce SD-FSMIS, a novel framework designed to effectively adapt the powerful pre-trained Stable Diffusion (SD) model for the FSMIS task. Our approach repurposes its conditional generative architecture by introducing two key components: a Support-Query Interaction (SQI) and a Visual-to-Textual Condition Translator (VTCT). Specifically, SQI provides a straightforward yet powerful means of adapting SD to the FSMIS paradigm. The VTCT module translates visual cues from the support set into an implicit textual embedding that guides the diffusion model, enabling precise conditioning of the generation process. Extensive experiments demonstrate that SD-FSMIS achieves competitive results compared to state-of-the-art methods in standard settings. Surprisingly, it also demonstrated excellent generalization ability in more challenging cross-domain scenarios. These findings highlight the immense potential of adapting large-scale generative models to advance data-efficient and robust medical image segmentation.

医学图像分割扩散模型少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。