arXiv:2602.02171cs.CV2026-02被引 1

用两阶段生成模型提升肺结节图像多样性与可控性

Lung Nodule Image Synthesis Driven by Two-Stage Generative Adversarial Networks

  • 分步生成:先建结构掩码,再转为CT图像
  • 在LUNA16上检测准确率提升4.6%,mAP提升4%
  • 适合医学图像增强和小样本检测研究者

肺结节CT数据集样本量有限且多样性不足,严重制约检测模型的性能与泛化能力。现有方法生成图像多样性差、可控性弱,存在纹理单一、解剖结构失真等问题。为此,我们提出两阶段生成对抗网络(TSGAN),通过解耦肺结节的形态结构与纹理特征,提升合成数据的多样性和空间可控性。第一阶段使用StyleGAN生成语义分割掩码图,编码肺结节与组织背景,控制解剖结构;第二阶段采用DL-Pix2Pix模型将掩码图转换为CT图像,引入局部重要性注意力捕捉局部特征,并利用动态权重多头窗注意力增强纹理与背景建模能力。实验表明,相较于原始数据集,TSGAN在LUNA16数据集上使检测准确率提升4.6%,平均精度均值(mAP)提升4%。结果证明,TSGAN可有效提升合成图像质量及检测模型性能。

原文摘要 · Abstract (English)

The limited sample size and insufficient diversity of lung nodule CT datasets severely restrict the performance and generalization ability of detection models. Existing methods generate images with insufficient diversity and controllability, suffering from issues such as monotonous texture features and distorted anatomical structures. Therefore, we propose a two-stage generative adversarial network (TSGAN) to enhance the diversity and spatial controllability of synthetic data by decoupling the morphological structure and texture features of lung nodules. In the first stage, StyleGAN is used to generate semantic segmentation mask images, encoding lung nodules and tissue backgrounds to control the anatomical structure of lung nodule images; The second stage uses the DL-Pix2Pix model to translate the mask map into CT images, employing local importance attention to capture local features, while utilizing dynamic weight multi-head window attention to enhance the modeling capability of lung nodule texture and background. Compared to the original dataset, the accuracy improved by 4.6% and mAP by 4% on the LUNA16 dataset. Experimental results demonstrate that TSGAN can enhance the quality of synthetic images and the performance of detection models.

图像生成医学影像生成对抗网络肺结节

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。