用疾病知识增强扩散模型,生成更真实准确的胸部X光片。
Diff-CXR: Report-to-CXR generation through a disease-knowledge enhanced diffusion model
- 通过隐空间噪声过滤和视觉感知文本学习,提升报告到X光片生成质量。
- 在MIMIC-CXR和IU-Xray数据集上FID和mAUC显著优于现有方法。
- 仅用1%真实数据+合成数据即可达到接近全量数据训练效果,适合临床应用。
文本到图像(TTI)生成在可控且多样化的图像生成中具有重要意义,尤其在医疗领域潜力巨大。尽管当前医学TTI方法在报告转胸部X光片(CXR)生成方面取得进展,但受限于医疗数据的固有特性,生成性能仍有瓶颈。本文提出一种新型疾病知识增强的扩散模型框架Diff-CXR,用于医学报告到CXR的生成。首先,设计潜空间噪声过滤策略,逐步学习异常的一般模式并在潜空间中去除噪声;其次,引入自适应视觉感知文本学习策略,在领域特定的视觉-语言模型中学习简洁关键的报告嵌入,为X光片生成提供文本引导;最后,通过精巧的控制适配器将通用疾病知识融入预训练TTI模型,构建疾病知识增强的扩散模型,实现逼真且精确的生成。实验表明,Diff-CXR在MIMIC-CXR和IU-Xray数据集上的FID和mAUC分别提升33.4%/8.0%和23.8%/56.4%,计算复杂度仅为29.641 GFLOPs。下游三组胸腔疾病分类任务及一项X光报告生成任务验证其有效性。值得注意的是,使用1%真实数据与合成数据联合训练的模型,可达到与全量数据训练相当的mAUC表现,展现出良好临床应用前景。
原文摘要 · Abstract (English)
Text-To-Image (TTI) generation is significant for controlled and diverse image generation with broad potential applications. Although current medical TTI methods have made some progress in report-to-Chest-Xray (CXR) generation, their generation performance may be limited due to the intrinsic characteristics of medical data. In this paper, we propose a novel disease-knowledge enhanced Diffusion-based TTI learning framework, named Diff-CXR, for medical report-to-CXR generation. First, to minimize the negative impacts of noisy data on generation, we devise a Latent Noise Filtering Strategy that gradually learns the general patterns of anomalies and removes them in the latent space. Then, an Adaptive Vision-Aware Textual Learning Strategy is designed to learn concise and important report embeddings in a domain-specific Vision-Language Model, providing textual guidance for Chest-Xray generation. Finally, by incorporating the general disease knowledge into the pretrained TTI model via a delicate control adapter, a disease-knowledge enhanced diffusion model is introduced to achieve realistic and precise report-to-CXR generation. Experimentally, our Diff-CXR outperforms previous SOTA medical TTI methods by 33.4\% / 8.0\% and 23.8\% / 56.4\% in the FID and mAUC score on MIMIC-CXR and IU-Xray, with the lowest computational complexity at 29.641 GFLOPs. Downstream experiments on three thorax disease classification benchmarks and one CXR-report generation benchmark demonstrate that Diff-CXR is effective in improving classical CXR analysis methods. Notably, models trained on the combination of 1\% real data and synthetic data can achieve a competitive mAUC score compared to models trained on all data, presenting promising clinical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。