用扩散模型生成高保真人造囊胚图像,解决辅助生殖数据少、不均衡难题。
Generating Synthetic Human Blastocyst Images for In-Vitro Fertilization Blastocyst Grading
- 基于扩散模型,按形态等级和焦距控制生成真实感囊胚图像。
- 合成数据可显著提升分类准确率,替换40%真实数据仍保持效果。
- 适合需增强数据的AI胚胎评估研究者与临床数据团队使用。
体外受精(IVF)的成功高度依赖于对第5天囊胚的形态学评估,但该过程主观性强且一致性差。尽管人工智能可实现标准化,但模型需要大量多样且平衡的数据,而实际中受限于数据稀缺、类别不平衡及隐私问题。现有生成式胚胎模型存在图像质量差、训练数据小、评估不稳健、缺乏临床相关性等问题。本文提出基于扩散的假囊胚影像生成框架(DIA),通过潜空间扩散模型生成高保真、新颖的第5天囊胚图像。模型支持基于Gardner分级和轴向焦深的细粒度控制。我们采用FID、记忆度量、胚胎学家图灵测试及三个下游分类任务进行严格评估。结果表明,生成图像与真实图像难以区分;更重要的是,用合成数据扩充不平衡数据集可显著提升分类准确率(p < 0.05)。即使在已有大而平衡的数据集上加入合成数据,性能仍显著提升,且在某些情况下,合成数据可替代高达40%的真实数据而不造成统计显著差异。DIA为解决胚胎数据稀缺与类别不平衡提供了可靠方案,通过生成可控、高质量的合成图像,可提升AI胚胎评估工具的性能、公平性与标准化水平。
原文摘要 · Abstract (English)
The success of in vitro fertilization (IVF) at many clinics relies on the accurate morphological assessment of day 5 blastocysts, a process that is often subjective and inconsistent. While artificial intelligence can help standardize this evaluation, models require large, diverse, and balanced datasets, which are often unavailable due to data scarcity, natural class imbalance, and privacy constraints. Existing generative embryo models can mitigate these issues but face several limitations, such as poor image quality, small training datasets, non-robust evaluation, and lack of clinically relevant image generation for effective data augmentation. Here, we present the Diffusion Based Imaging Model for Artificial Blastocysts (DIA) framework, a set of latent diffusion models trained to generate high-fidelity, novel day 5 blastocyst images. Our models provide granular control by conditioning on Gardner-based morphological categories and z-axis focal depth. We rigorously evaluated the models using FID, a memorization metric, an embryologist Turing test, and three downstream classification tasks. Our results show that DIA models generate realistic images that embryologists could not reliably distinguish from real images. Most importantly, we demonstrated clear clinical value. Augmenting an imbalanced dataset with synthetic images significantly improved classification accuracy (p < 0.05). Also, adding synthetic images to an already large, balanced dataset yielded statistically significant performance gains, and synthetic data could replace up to 40% of real data in some cases without a statistically significant loss in accuracy. DIA provides a robust solution for mitigating data scarcity and class imbalance in embryo datasets. By generating novel, high-fidelity, and controllable synthetic images, our models can improve the performance, fairness, and standardization of AI embryo assessment tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。