arXiv:2508.17844cs.CVcs.LG2025-08ICCV被引 6

用文本引导的扩散模型生成医学图像异常,提升罕见病灶分割效果。

Diffusion-Based Data Augmentation for Medical Image Segmentation

  • 基于文本和掩码的扩散模型,在正常图像中合成病灶。
  • 通过潜在空间分割网络动态验证生成质量,提升定位准确率。
  • 适用于小息肉、平坦病灶等难检测病灶,适合医疗筛查场景。

医学图像分割模型因病理标注数据稀少,难以识别罕见异常。本文提出 DiffAug 框架,结合文本引导的扩散生成与自动分割验证,解决该问题。方法利用条件于医学文本描述和空间掩码的潜空间扩散模型,在正常图像上通过修复生成异常。生成样本经由潜在空间分割网络进行动态质量评估,确保定位准确并支持单步推理。文本提示来自医学文献,无需人工标注即可生成多样化异常类型。验证机制依据空间精度过滤合成样本,实现高效高质量生成。在 CVC-ClinicDB、Kvasir-SEG、REFUGE2 三个医学影像基准上评估,相比基线模型,Dice 分数提升 8-10%,对小息肉和扁平病灶等挑战性病例,假阴性率降低最高达 28%,显著提升早期筛查性能。

原文摘要 · Abstract (English)

Medical image segmentation models struggle with rare abnormalities due to scarce annotated pathological data. We propose DiffAug a novel framework that combines textguided diffusion-based generation with automatic segmentation validation to address this challenge. Our proposed approach uses latent diffusion models conditioned on medical text descriptions and spatial masks to synthesize abnormalities via inpainting on normal images. Generated samples undergo dynamic quality validation through a latentspace segmentation network that ensures accurate localization while enabling single-step inference. The text prompts, derived from medical literature, guide the generation of diverse abnormality types without requiring manual annotation. Our validation mechanism filters synthetic samples based on spatial accuracy, maintaining quality while operating efficiently through direct latent estimation. Evaluated on three medical imaging benchmarks (CVC-ClinicDB, Kvasir-SEG, REFUGE2), our framework achieves state-of-the-art performance with 8-10% Dice improvements over baselines and reduces false negative rates by up to 28% for challenging cases like small polyps and flat lesions critical for early detection in screening applications.

医学图像扩散模型数据增强分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。