用GAN生成肺部X光片提升肺炎检测,但效果需谨慎评估
A Comparative Study of GAN-Based Deep Learning Models for Pneumonia Detection in Chest X-Rays

- 对比VGG19等4种模型,用GAN合成数据增强训练
- MobileNetV2达88%准确率,自定义CNN召回率达92.67%
- 合成数据提升训练准确率但未改善验证性能,需关注数据质量
本研究评估了VGG19、MobileNetV2、ResNet50及自定义CNN在胸部X光片中肺炎分类的表现,并探索基于生成对抗网络(GAN)的合成数据增强效果。MobileNetV2实现最高准确率88%,且类别平衡性良好。自定义CNN的肺炎召回率为92.67%,精确率为79.43%,体现精确率与召回率的权衡。通过准确率、F1分数、精确率、召回率、混淆矩阵和训练曲线评估模型性能。将合成肺炎图像与真实图像结合,探究增强对分类性能的影响。在VGG19对比实验中,使用增强数据训练的准确率接近100%,但验证准确率仅约50%,低于真实数据训练的验证表现。该结果表明,当前合成数据增强未带来验证性能提升,提示需进一步评估合成图像质量与训练设置。
原文摘要 · Abstract (English)
This study evaluates pneumonia classification in chest X-rays using VGG19, MobileNetV2, ResNet50, and a custom CNN, and explores Generative Adversarial Network (GAN)-based synthetic data augmentation. MobileNetV2 achieved the highest reported accuracy of 88% with balanced class-wise performance. The custom CNN achieved pneumonia recall of 92.67% and precision of 79.43%, highlighting a precision-recall trade-off. Accuracy, F1-score, precision, recall, confusion matrices, and training curves were used to assess performance. Synthetic pneumonia images were combined with real images to investigate whether augmentation could improve classification performance. In the reported VGG19 comparison, augmented-data training accuracy reached approximately 100%, while validation accuracy remained near 50%, below the real-data validation accuracy. This experiment therefore did not demonstrate a validation-performance benefit from GAN augmentation. The classifier comparison highlights differences in accuracy and pneumonia recall, while the augmentation experiment indicates the need for further evaluation of synthetic-image quality and training settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。