用视频扩散模型实现长距离可控场景生成,支持精确姿态控制。
StarGen: A Spatiotemporal Autoregression Framework with Video Diffusion Model for Scalable and Controllable Scene Generation
- 基于视频扩散模型的自回归框架,逐段生成视频片段。
- 生成时融合前后帧与邻近视角图像,提升时空一致性。
- 适合稀疏视图插值、持续视角生成等任务,可精确控制相机姿态。
近期大型重建与生成模型显著提升了场景重建与新视角生成能力。然而,受计算资源限制,这些大模型每次推理仅能处理小范围内容,难以实现长距离一致的场景生成。为此,我们提出StarGen,一种新颖框架:将预训练视频扩散模型以自回归方式使用,实现长距离场景生成。每个视频片段的生成均依赖于空间相邻图像的3D投影以及先前生成片段的时间重叠图像,从而在长序列中保持良好的时空一致性,并实现精确的姿态控制。该时空条件可适配多种输入,支持稀疏视图插值、持续视角生成和布局约束的城市生成等多种任务。定量与定性评估表明,相比当前最优方法,StarGen在可扩展性、保真度和姿态准确性方面均有显著提升。
原文摘要 · Abstract (English)
Recent advances in large reconstruction and generative models have significantly improved scene reconstruction and novel view generation. However, due to compute limitations, each inference with these large models is confined to a small area, making long-range consistent scene generation challenging. To address this, we propose StarGen, a novel framework that employs a pre-trained video diffusion model in an autoregressive manner for long-range scene generation. The generation of each video clip is conditioned on the 3D warping of spatially adjacent images and the temporally overlapping image from previously generated clips, improving spatiotemporal consistency in long-range scene generation with precise pose control. The spatiotemporal condition is compatible with various input conditions, facilitating diverse tasks, including sparse view interpolation, perpetual view generation, and layout-conditioned city generation. Quantitative and qualitative evaluations demonstrate StarGen's superior scalability, fidelity, and pose accuracy compared to state-of-the-art methods. Project page: https://zju3dv.github.io/StarGen.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。