用文字生成眼底荧光血管造影动态视频,辅助医生学习诊断。
FFA Sora, video generation as fundus fluorescein angiography simulator
- 通过小波流变分自编码器与扩散Transformer,将报告转为动态视频。
- 生成视频与文本匹配度高,客观指标显示FVD=329.78,VQAScore=0.61。
- 保护患者隐私且视觉质量好,适合医学教育与临床培训使用。
眼底荧光血管造影(FFA)对视网膜血管病诊断至关重要,但初学者常难以解读图像。本研究开发了FFA Sora,一种文生视频模型,利用小波流变分自编码器(WF-VAE)和扩散Transformer(DiT),将FFA报告转化为动态视频。模型在匿名数据集上训练,客观评估显示:弗雷切特视频距离(FVD)为329.78,感知图像块相似度(LPIPS)为0.48,视觉问答得分(VQAScore)为0.61。特定评估表明生成视频与文本提示具有可接受的对齐性,BERTScore为0.35。检索评估中,模型表现出良好的隐私保护能力,平均召回率@K为0.073。人工评估显示视频视觉质量良好,平均得分1.570(1为最佳,5为最差)。该模型解决了大规模FFA数据共享带来的隐私问题,并有助于医学教育。
原文摘要 · Abstract (English)
Fundus fluorescein angiography (FFA) is critical for diagnosing retinal vascular diseases, but beginners often struggle with image interpretation. This study develops FFA Sora, a text-to-video model that converts FFA reports into dynamic videos via a Wavelet-Flow Variational Autoencoder (WF-VAE) and a diffusion transformer (DiT). Trained on an anonymized dataset, FFA Sora accurately simulates disease features from the input text, as confirmed by objective metrics: Frechet Video Distance (FVD) = 329.78, Learned Perceptual Image Patch Similarity (LPIPS) = 0.48, and Visual-question-answering Score (VQAScore) = 0.61. Specific evaluations showed acceptable alignment between the generated videos and textual prompts, with BERTScore of 0.35. Additionally, the model demonstrated strong privacy-preserving performance in retrieval evaluations, achieving an average Recall@K of 0.073. Human assessments indicated satisfactory visual quality, with an average score of 1.570(scale: 1 = best, 5 = worst). This model addresses privacy concerns associated with sharing large-scale FFA data and enhances medical education.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。