arXiv:2601.09697cs.CV2026-01被引 2

用稀疏关键帧+3D渲染,40倍提速视频生成

Efficient Camera-Controlled Video Generation of Static Scenes via Sparse Diffusion and 3D Rendering

  • 先生成少量关键帧,再通过3D重建渲染补全视频
  • 20秒视频生成速度提升40倍以上,保持画面稳定清晰
  • 自动调节关键帧数量,适配复杂或简单镜头运动

基于扩散模型的现代视频生成方法虽能产出高度逼真的视频片段,但计算效率极低,生成几秒视频需数分钟GPU时间。这严重阻碍了其在需要实时交互的应用(如具身AI、VR/AR)中的部署。本文提出一种相机控制的静态场景视频生成新策略:利用扩散模型生成稀疏的关键帧,再通过3D重建与渲染合成完整视频。将关键帧提升至3D表示并渲染中间视角,使生成成本分摊到数百帧,同时保证几何一致性。进一步提出一个模型,可预测给定相机轨迹所需的最优关键帧数量,实现计算资源的自适应分配。最终方法SRENDER对简单轨迹使用极稀疏关键帧,复杂运动则采用更密集采样。该方法在生成20秒视频时,速度比纯扩散基线快40倍以上,同时保持高视觉保真度和时间稳定性,为高效可控的视频合成提供了可行路径。

原文摘要 · Abstract (English)

Modern video generative models based on diffusion models can produce very realistic clips, but they are computationally inefficient, often requiring minutes of GPU time for just a few seconds of video. This inefficiency poses a critical barrier to deploying generative video in applications that require real-time interactions, such as embodied AI and VR/AR. This paper explores a new strategy for camera-conditioned video generation of static scenes: using diffusion-based generative models to generate a sparse set of keyframes, and then synthesizing the full video through 3D reconstruction and rendering. By lifting keyframes into a 3D representation and rendering intermediate views, our approach amortizes the generation cost across hundreds of frames while enforcing geometric consistency. We further introduce a model that predicts the optimal number of keyframes for a given camera trajectory, allowing the system to adaptively allocate computation. Our final method, SRENDER, uses very sparse keyframes for simple trajectories and denser ones for complex camera motion. This results in video generation that is more than 40 times faster than the diffusion-based baseline in generating 20 seconds of video, while maintaining high visual fidelity and temporal stability, offering a practical path toward efficient and controllable video synthesis.

视频生成扩散模型3D重建高效生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。