用96万张胸片训练生成模型,可合成逼真医学影像提升诊断准确率。
A Generative Foundation Model for Chest Radiography
- 基于扩散变换器架构,支持文本/掩码/框引导生成胸片。
- 在小样本下提升病种分类、检测与分割性能,优于传统方法。
- 能生成多样化患者数据,帮助发现并减少性别年龄偏见。
医疗影像标注数据稀缺严重制约可靠AI模型发展。我们提出ChexGen,一种用于胸部X光的生成式视觉-语言基础模型,实现文本、掩码和边界框引导的统一合成框架。该模型基于潜变分扩散变换器架构,在迄今最大的精选胸片数据集(96万对影像-报告)上预训练。经专家评估与量化指标验证,其合成图像精度高。实验表明,利用ChexGen进行数据增强或监督预训练,仅需少量真实数据即可显著提升疾病分类、检测与分割任务表现。此外,该模型可生成多样化患者群体,有效识别并缓解人口统计学偏差,推动更公平的医疗AI系统建设。
原文摘要 · Abstract (English)
The scarcity of well-annotated diverse medical images is a major hurdle for developing reliable AI models in healthcare. Substantial technical advances have been made in generative foundation models for natural images. Here we develop `ChexGen', a generative vision-language foundation model that introduces a unified framework for text-, mask-, and bounding box-guided synthesis of chest radiographs. Built upon the latent diffusion transformer architecture, ChexGen was pretrained on the largest curated chest X-ray dataset to date, consisting of 960,000 radiograph-report pairs. ChexGen achieves accurate synthesis of radiographs through expert evaluations and quantitative metrics. We demonstrate the utility of ChexGen for training data augmentation and supervised pretraining, which led to performance improvements across disease classification, detection, and segmentation tasks using a small fraction of training data. Further, our model enables the creation of diverse patient cohorts that enhance model fairness by detecting and mitigating demographic biases. Our study supports the transformative role of generative foundation models in building more accurate, data-efficient, and equitable medical AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。