arXiv:2503.07209cs.CVcs.LG2025-03被引 13

用扩散模型自动生成肺部X光图的精确语义掩码

Synthetic Lung X-ray Generation through Cross-Attention and Affinity Transformation

  • 通过文本-图像交叉注意力,将文本生成扩展到语义掩码生成
  • 合成数据训练的分割模型性能媲美甚至超越真实数据训练模型
  • 适合医疗影像数据稀缺场景下的研究者与工程师使用

医学影像的收集与标注耗时且资源密集。通过扩散模型生成合成数据可降低成本。本文提出一种基于文本-图像对训练的稳定扩散模型,自动从合成肺部X光图生成准确的语义掩码。该方法利用文本与图像间的交叉注意力映射,将文本驱动的图像生成拓展至掩码生成;通过文本引导的交叉注意力信息定位图像特定区域,并结合创新技术生成高分辨率、类别区分的像素级掩码。实验表明,使用该方法生成的合成数据训练的分割模型,性能与真实数据训练模型相当,部分情况下更优,验证了方法的有效性及其在医学影像分析中的变革潜力。

原文摘要 · Abstract (English)

Collecting and annotating medical images is a time-consuming and resource-intensive task. However, generating synthetic data through models such as Diffusion offers a cost-effective alternative. This paper introduces a new method for the automatic generation of accurate semantic masks from synthetic lung X-ray images based on a stable diffusion model trained on text-image pairs. This method uses cross-attention mapping between text and image to extend text-driven image synthesis to semantic mask generation. It employs text-guided cross-attention information to identify specific areas in an image and combines this with innovative techniques to produce high-resolution, class-differentiated pixel masks. This approach significantly reduces the costs associated with data collection and annotation. The experimental results demonstrate that segmentation models trained on synthetic data generated using the method are comparable to, and in some cases even better than, models trained on real datasets. This shows the effectiveness of the method and its potential to revolutionize medical image analysis.

医学影像扩散模型语义分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。