用一张眼底照片生成动态血管造影视频,非侵入式诊断新方案
Fundus to Fluorescein Angiography Video Generation as a Retinal Generative Foundation Model
- 基于单张眼底照片生成动态荧光造影视频,采用自回归生成对抗网络
- 视频质量达FVD 1497.12、PSNR 11.77,临床专家验证生成视频真实性
- 模型可零样本/少样本迁移至10个外部数据集,适用于多种眼科任务
眼底荧光血管造影(FFA)对视网膜血管病变的诊断与监测至关重要,但其侵入性及获取难度高于彩色眼底(CF)成像。现有方法仅能将CF图像转为静态FFA,忽略病灶动态变化。我们提出Fundus2Video,一种自回归生成对抗网络(GAN),可从单张CF图像生成动态FFA视频。该模型在视频生成上表现优异,实现FVD 1497.12和PSNR 11.77,临床专家验证了生成视频的保真度。此外,其生成器在10个外部公共数据集上展现卓越下游迁移能力,涵盖血管分割、视网膜疾病诊断、系统性疾病预测及多模态检索任务,具备出色的零样本与少样本性能。这些成果表明Fundus2Video是替代侵入性FFA检查的强大非侵入方案,并作为捕捉静态与时间特征的视网膜生成基础模型,实现复杂跨模态关系建模。
原文摘要 · Abstract (English)
Fundus fluorescein angiography (FFA) is crucial for diagnosing and monitoring retinal vascular issues but is limited by its invasive nature and restricted accessibility compared to color fundus (CF) imaging. Existing methods that convert CF images to FFA are confined to static image generation, missing the dynamic lesional changes. We introduce Fundus2Video, an autoregressive generative adversarial network (GAN) model that generates dynamic FFA videos from single CF images. Fundus2Video excels in video generation, achieving an FVD of 1497.12 and a PSNR of 11.77. Clinical experts have validated the fidelity of the generated videos. Additionally, the model's generator demonstrates remarkable downstream transferability across ten external public datasets, including blood vessel segmentation, retinal disease diagnosis, systemic disease prediction, and multimodal retrieval, showcasing impressive zero-shot and few-shot capabilities. These findings position Fundus2Video as a powerful, non-invasive alternative to FFA exams and a versatile retinal generative foundation model that captures both static and temporal retinal features, enabling the representation of complex inter-modality relationships.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。