arXiv:2507.18214cs.CV2025-07中稿 · MICCAI 2025

用轻量蒸馏提升扩散模型医学分割性能,不增加推理开销

LEAF: Latent Diffusion with Efficient Encoder Distillation for Aligned Features in Medical Image Segmentation

  • 直接预测分割图替代噪声预测,降低结果方差
  • 蒸馏对齐卷积层与视觉变压器特征,提升语义一致性
  • 无需改架构、不增参数,适合临床部署

利用扩散模型的强大能力在医学图像分割任务中已取得显著成效。然而,现有方法通常直接迁移原始训练过程,未针对分割任务进行适配;且常用预训练扩散模型在特征提取方面仍存在不足。为此,我们提出基于潜在扩散模型的医学图像分割方法LEAF。在微调过程中,将原有的噪声预测方式替换为直接预测分割图,从而降低分割结果的方差。同时,采用特征蒸馏方法,将卷积层的隐藏状态与基于Transformer的视觉编码器特征对齐。实验表明,该方法在多个不同疾病类型的分割数据集上均提升了原扩散模型的性能。值得注意的是,本方法不改变模型架构,推理阶段不增加参数或计算量,具有高效性。

原文摘要 · Abstract (English)

Leveraging the powerful capabilities of diffusion models has yielded quite effective results in medical image segmentation tasks. However, existing methods typically transfer the original training process directly without specific adjustments for segmentation tasks. Furthermore, the commonly used pre-trained diffusion models still have deficiencies in feature extraction. Based on these considerations, we propose LEAF, a medical image segmentation model grounded in latent diffusion models. During the fine-tuning process, we replace the original noise prediction pattern with a direct prediction of the segmentation map, thereby reducing the variance of segmentation results. We also employ a feature distillation method to align the hidden states of the convolutional layers with the features from a transformer-based vision encoder. Experimental results demonstrate that our method enhances the performance of the original diffusion model across multiple segmentation datasets for different disease types. Notably, our approach does not alter the model architecture, nor does it increase the number of parameters or computation during the inference phase, making it highly efficient.

医学分割扩散模型特征对齐轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。