arXiv:2507.15361eess.IVcs.AI2025-07中稿 · CVGMMI Workshop at…

用文本控制生成真实肠镜图像,提升小样本分割效果

Latent Space Synergy: Text-Guided Data Augmentation for Direct Diffusion Biomedical Segmentation

  • 用文本引导的扩散模型生成逼真肠息肉图像,增强训练数据
  • 单步推理速度提升T倍,在CVC-ClinicDB上达96.0% Dice
  • 适合医疗资源有限场景,兼顾精度与实时性

医学图像分割面临数据稀缺问题,尤其在肠息肉检测中,标注需专业经验。本文提出SynDiff框架,结合文本引导的合成数据生成与基于扩散模型的高效分割。该方法利用潜空间扩散模型,通过文本条件修复生成临床上真实的合成肠息肉,以语义多样样本扩充有限训练数据。不同于传统需迭代去噪的扩散方法,我们引入直接潜变量估计,实现单步推断,获得T倍计算加速。在CVC-ClinicDB数据集上,SynDiff达到96.0% Dice和92.9% IoU,同时保持实时性能,适用于临床部署。结果表明,受控的合成数据增强可提升分割鲁棒性且不引发分布偏移。SynDiff弥合了数据密集型深度学习与临床约束之间的差距,为资源受限医疗环境提供高效解决方案。

原文摘要 · Abstract (English)

Medical image segmentation suffers from data scarcity, particularly in polyp detection where annotation requires specialized expertise. We present SynDiff, a framework combining text-guided synthetic data generation with efficient diffusion-based segmentation. Our approach employs latent diffusion models to generate clinically realistic synthetic polyps through text-conditioned inpainting, augmenting limited training data with semantically diverse samples. Unlike traditional diffusion methods requiring iterative denoising, we introduce direct latent estimation enabling single-step inference with T x computational speedup. On CVC-ClinicDB, SynDiff achieves 96.0% Dice and 92.9% IoU while maintaining real-time capability suitable for clinical deployment. The framework demonstrates that controlled synthetic augmentation improves segmentation robustness without distribution shift. SynDiff bridges the gap between data-hungry deep learning models and clinical constraints, offering an efficient solution for deployment in resourcelimited medical settings.

医学分割扩散模型数据增强实时推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。