用静态超声图生成逼真动态视频,解决数据少难题
Ultrasound Image-to-Video Synthesis via Latent Dynamic Diffusion Models
- 基于潜在动态扩散模型,将单张图像转为连贯视频序列
- 在BUSV数据集上合成视频质量高,分类性能提升显著
- 适合超声视频分析、数据增强研究者使用
超声视频分类有助于自动化诊断,是当前重要研究方向。然而公开可用的超声视频数据集仍十分稀缺,制约了有效视频分类模型的发展。为此,我们提出从大量易获取的静态超声图像中合成合理真实的超声视频,以缓解数据不足问题。我们引入潜动态扩散模型(LDDM),高效地将静态图像转换为具有真实视频特性的动态序列。在BUSV基准上,我们展示了优异的定量结果和视觉上令人信服的合成视频。值得注意的是,使用真实数据与LDDM合成数据混合训练分类模型,性能明显优于仅使用真实数据,表明该方法成功模拟了对区分任务至关重要的动态特性。本图像到视频的方法为推进超声视频分析提供了有效的数据增强方案。代码已开源:https://github.com/MedAITech/U_I2V。
原文摘要 · Abstract (English)
Ultrasound video classification enables automated diagnosis and has emerged as an important research area. However, publicly available ultrasound video datasets remain scarce, hindering progress in developing effective video classification models. We propose addressing this shortage by synthesizing plausible ultrasound videos from readily available, abundant ultrasound images. To this end, we introduce a latent dynamic diffusion model (LDDM) to efficiently translate static images to dynamic sequences with realistic video characteristics. We demonstrate strong quantitative results and visually appealing synthesized videos on the BUSV benchmark. Notably, training video classification models on combinations of real and LDDM-synthesized videos substantially improves performance over using real data alone, indicating our method successfully emulates dynamics critical for discrimination. Our image-to-video approach provides an effective data augmentation solution to advance ultrasound video analysis. Code is available at https://github.com/MedAITech/U_I2V.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。