用合成数据训练医疗视觉语言模型,效果竟优于真实数据。
Can Medical Vision-Language Pre-training Succeed with Purely Synthetic Data?

- 用生成模型自动构建高质量合成医学影像-文本对
- 纯合成数据训练使零样本分类AUC提升3.8%
- 适合想低成本训练医疗AI的研究者参考
医疗视觉语言预训练(MedVLP)在医学图像理解的零样本任务中取得显著进展,但通常需要大规模、高质量的配对图像-文本数据,而这类数据在医疗领域稀缺。近年来,大语言模型和扩散模型的发展使得生成大规模合成图像-文本对成为可能。本研究探讨:能否仅用合成数据实现成功训练?我们采用现成的生成模型创建合成放射科报告与配对的胸部X光片,并提出自动化流程构建多样且高质量的合成数据集,从而严格控制变量,聚焦数据本身的影响。结果表明,仅使用合成数据训练的MedVLP模型在零样本分类上的平均AUC比真实数据训练高出3.8%;结合合成与真实数据进一步提升9.07%。此外,在零样本定位、微调分类与分割任务中,合成或混合数据训练的模型均优于纯真实数据训练模型。分析显示,设计良好的合成数据可超越受限于低质量样本和长尾分布的真实数据集。
原文摘要 · Abstract (English)
Medical Vision-Language Pre-training (MedVLP) has made significant progress in enabling zero-shot tasks for medical image understanding. However, training MedVLP models typically requires large-scale datasets with paired, high-quality image-text data, which are scarce in the medical domain. Recent advancements in Large Language Models (LLMs) and diffusion models have made it possible to generate large-scale synthetic image-text pairs. This raises the question: "Can MedVLP succeed using purely synthetic data?" To address this, we use off-the-shelf generative models to create synthetic radiology reports and paired Chest X-ray (CXR) images, and propose an automated pipeline to build a diverse, high-quality synthetic dataset, enabling a rigorous study that isolates model and training settings, focusing entirely from the data perspective. Our results show that MedVLP models trained exclusively on synthetic data outperform those trained on real data by 3.8% in averaged AUC on zero-shot classification. Moreover, using a combination of synthetic and real data leads to a further improvement of 9.07%. Additionally, MedVLP models trained on synthetic or mixed data consistently outperform those trained on real data in zero-shot grounding, as well as in fine-tuned classification and segmentation tasks. Our analysis suggests MedVLP trained on well-designed synthetic data can outperform models trained on real datasets, which may be limited by low-quality samples and long-tailed distributions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。