用图像与视频扩散模型联合生成,提升画质与时间连贯性。
JVID: Joint Video-Image Diffusion for Visual-Quality and Temporal-Consistency in Video Generation

- 融合图像与视频扩散模型,在反向生成中协同优化
- 生成视频在画质和时序一致性上均有显著提升
- 适合需要高质量、连贯视频生成的研究与应用
我们提出联合视频-图像扩散模型(JVID),一种生成高质量且时序一致视频的新方法。通过整合在图像数据上训练的潜在图像扩散模型(LIDM)和在视频数据上训练的潜在视频扩散模型(LVDM),在反向扩散过程中联合使用两者:LIDM提升画面质量,LVDM保障时序连贯性。该机制有效处理了视频生成中的复杂时空动态。实验结果表明,该方法在定性和定量层面均实现了更真实、更连贯的视频生成效果。
原文摘要 · Abstract (English)
We introduce the Joint Video-Image Diffusion model (JVID), a novel approach to generating high-quality and temporally coherent videos. We achieve this by integrating two diffusion models: a Latent Image Diffusion Model (LIDM) trained on images and a Latent Video Diffusion Model (LVDM) trained on video data. Our method combines these models in the reverse diffusion process, where the LIDM enhances image quality and the LVDM ensures temporal consistency. This unique combination allows us to effectively handle the complex spatio-temporal dynamics in video generation. Our results demonstrate quantitative and qualitative improvements in producing realistic and coherent videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。