arXiv:2411.10004eess.IVcs.AI2024-11被引 9

用文本生成眼底图像,提升罕见眼病诊断准确率

EyeDiff: text-to-image diffusion model improves rare eye disease diagnosis

  • 基于文本生成多模态眼底图像,模拟真实病变特征
  • 生成数据使罕见病识别准确率超越传统数据扩充方法
  • 适合眼科疾病模型训练、医疗数据稀缺场景

视力威胁性视网膜疾病的发病率上升,给全球医疗系统带来巨大负担。深度学习虽可实现自动筛查,但需大量标注数据,而罕见眼病的多模态眼科影像采集与标注面临现实挑战。本文提出EyeDiff,一种文本到图像的生成模型,可从自然语言提示生成多模态眼科图像,并评估其在常见与罕见疾病诊断中的适用性。EyeDiff在八个大规模数据集上训练,采用先进的潜在扩散模型,覆盖14种眼科影像模态和80余种眼病,并适配10个跨国外部数据集。生成图像能准确捕捉关键病灶特征,在客观指标与专家评估中均与文本提示高度一致。整合生成图像显著提升少数类及罕见眼病的检测准确率,优于传统过采样方法,有效缓解罕见病数据不平衡与数据不足问题,为构建眼科专家级诊断模型提供变革性解决方案。

原文摘要 · Abstract (English)

The rising prevalence of vision-threatening retinal diseases poses a significant burden on the global healthcare systems. Deep learning (DL) offers a promising solution for automatic disease screening but demands substantial data. Collecting and labeling large volumes of ophthalmic images across various modalities encounters several real-world challenges, especially for rare diseases. Here, we introduce EyeDiff, a text-to-image model designed to generate multimodal ophthalmic images from natural language prompts and evaluate its applicability in diagnosing common and rare diseases. EyeDiff is trained on eight large-scale datasets using the advanced latent diffusion model, covering 14 ophthalmic image modalities and over 80 ocular diseases, and is adapted to ten multi-country external datasets. The generated images accurately capture essential lesional characteristics, achieving high alignment with text prompts as evaluated by objective metrics and human experts. Furthermore, integrating generated images significantly enhances the accuracy of detecting minority classes and rare eye diseases, surpassing traditional oversampling methods in addressing data imbalance. EyeDiff effectively tackles the issue of data imbalance and insufficiency typically encountered in rare diseases and addresses the challenges of collecting large-scale annotated images, offering a transformative solution to enhance the development of expert-level diseases diagnosis models in ophthalmic field.

文本生成眼病诊断数据增强扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。