无需训练,用单视频生成多视角4D视频。
Zero4D: Training-Free 4D Video Generation From Single Video Using Off-the-Shelf Video Diffusion
- 用现成视频扩散模型,通过关键帧合成与插值生成多视角视频。
- 在无额外训练下实现空间时间一致性,支持新视角相机轨迹。
- 适合快速生成多视角视频的开发者和内容创作者。
多视角或4D视频生成已成为重要研究方向。然而,现有方法仍受限于依赖多个视频扩散模型或需大量计算资源训练完整4D扩散模型,且真实世界4D数据稀缺。为此,我们提出首个无需训练的4D视频生成方法,仅使用现成视频扩散模型,从单个输入视频生成多视角视频。该方法包含两步:(1)将时空采样网格中的边缘帧设为关键帧,利用基于深度的扭曲技术引导,通过视频扩散模型合成,确保生成帧间结构一致性,保持空间与时间连贯性;(2)再用视频扩散模型对剩余帧进行插值,构建完整且时间连贯的采样网格,维持空间与时间一致性。本方法将单个视频扩展为沿新相机轨迹的多视角视频,实现无训练、完全利用现成模型的高效多视角生成。
原文摘要 · Abstract (English)
Multi-view or 4D video generation has emerged as a significant research topic. Nonetheless, recent approaches to 4D generation still struggle with fundamental limitations, as they primarily rely on harnessing multiple video diffusion models with additional training or compute-intensive training of a full 4D diffusion model with limited real-world 4D data and large computational costs. To address these challenges, here we propose the first training-free 4D video generation method that leverages the off-the-shelf video diffusion models to generate multi-view videos from a single input video. Our approach consists of two key steps: (1) By designating the edge frames in the spatio-temporal sampling grid as key frames, we first synthesize them using a video diffusion model, leveraging a depth-based warping technique for guidance. This approach ensures structural consistency across the generated frames, preserving spatial and temporal coherence. (2) We then interpolate the remaining frames using a video diffusion model, constructing a fully populated and temporally coherent sampling grid while preserving spatial and temporal consistency. Through this approach, we extend a single video into a multi-view video along novel camera trajectories while maintaining spatio-temporal consistency. Our method is training-free and fully utilizes an off-the-shelf video diffusion model, offering a practical and effective solution for multi-view video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。