对比GAN与扩散模型生成胸片的临床可用性,发现两者各有优劣。
Perceptual Evaluation of GANs and Diffusion Models for Generating X-rays
- 用放射科医生盲评真实与合成胸片,评估生成质量
- 扩散模型整体更逼真,但GAN在特定异常检测上更准
- 揭示医生识别假片的视觉线索,指导模型改进
生成式图像模型在自然与医学影像领域均取得显著进展。在医疗场景中,这类技术可缓解数据稀缺问题,尤其对低发病率异常(如肺不张、肺部浸润、胸腔积液、心脏轮廓增大)具有重要意义,能提升AI诊断与分割模型性能。然而,合成图像的真实性和临床实用性仍存疑,因生成质量不佳会削弱模型泛化能力与可信度。本研究评估了当前最先进的生成对抗网络(GANs)与扩散模型(DMs)在生成四类胸部异常(肺不张/AT、肺部浸润/LO、胸腔积液/PE、心脏轮廓增大/ECS)相关胸片上的表现。基于MIMIC-CXR真实数据集与两类模型生成的合成图像,我们邀请三位经验各异的放射科医生参与读者研究,判断图像真伪并评估视觉特征与目标异常的一致性。结果显示,尽管扩散模型生成图像整体更逼真,但生成对抗网络在特定条件(如无心脏轮廓增大)下的识别准确率更高。我们进一步识别出放射科医生用于辨别合成图像的关键视觉线索,揭示了当前模型在感知层面的差距。这些发现凸显了两类模型的互补优势,并指出需进一步优化以确保生成模型可可靠扩充人工智能诊断系统的训练数据。
原文摘要 · Abstract (English)
Generative image models have achieved remarkable progress in both natural and medical imaging. In the medical context, these techniques offer a potential solution to data scarcity-especially for low-prevalence anomalies that impair the performance of AI-driven diagnostic and segmentation tools. However, questions remain regarding the fidelity and clinical utility of synthetic images, since poor generation quality can undermine model generalizability and trust. In this study, we evaluate the effectiveness of state-of-the-art generative models-Generative Adversarial Networks (GANs) and Diffusion Models (DMs)-for synthesizing chest X-rays conditioned on four abnormalities: Atelectasis (AT), Lung Opacity (LO), Pleural Effusion (PE), and Enlarged Cardiac Silhouette (ECS). Using a benchmark composed of real images from the MIMIC-CXR dataset and synthetic images from both GANs and DMs, we conducted a reader study with three radiologists of varied experience. Participants were asked to distinguish real from synthetic images and assess the consistency between visual features and the target abnormality. Our results show that while DMs generate more visually realistic images overall, GANs can report better accuracy for specific conditions, such as absence of ECS. We further identify visual cues radiologists use to detect synthetic images, offering insights into the perceptual gaps in current models. These findings underscore the complementary strengths of GANs and DMs and point to the need for further refinement to ensure generative models can reliably augment training datasets for AI diagnostic systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。