用扩散模型生成任意视角的新画面,支持长视频无缝循环。
Stable Virtual Camera: Generative View Synthesis with Diffusion Models

- 基于扩散模型设计,无需3D重建即可生成新视角。
- 可生成长达30秒的高质量连续视频,且首尾自然衔接。
- 通用性强,适配多种场景和视角变换,部署简单。
我们提出Stable Virtual Camera(Seva),一种通用扩散模型,可基于任意数量输入视角和目标相机位置生成新视图。现有方法在大幅视角变化或时间上连续生成方面表现不佳,且依赖特定任务配置。本方法通过简洁的模型设计、优化的训练策略和灵活采样机制,在测试时实现跨任务泛化,生成结果保持高一致性,无需额外基于3D表示的蒸馏过程,从而简化了真实场景下的视图合成。此外,我们的方法可生成长达半分钟的高质量视频,并实现无缝循环闭合。大量基准测试表明,Seva在不同数据集和设置下均优于现有方法。
原文摘要 · Abstract (English)
We present Stable Virtual Camera (Seva), a generalist diffusion model that creates novel views of a scene, given any number of input views and target cameras. Existing works struggle to generate either large viewpoint changes or temporally smooth samples, while relying on specific task configurations. Our approach overcomes these limitations through simple model design, optimized training recipe, and flexible sampling strategy that generalize across view synthesis tasks at test time. As a result, our samples maintain high consistency without requiring additional 3D representation-based distillation, thus streamlining view synthesis in the wild. Furthermore, we show that our method can generate high-quality videos lasting up to half a minute with seamless loop closure. Extensive benchmarking demonstrates that Seva outperforms existing methods across different datasets and settings. Project page with code and model: https://stable-virtual-camera.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。