用单张或稀疏图生成高质量新视角视频,靠扩散模型和点云先验。
ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis

- 结合扩散模型生成能力与点表示的粗略3D信息,控制相机姿态生成视频帧。
- 通过迭代合成与轨迹规划,扩展可生成的新视角范围和覆盖区域。
- 适合需要高保真新视角生成的场景,如实时沉浸式体验和文本生成3D内容。
尽管神经3D重建技术取得进展,但对密集多视角采集的依赖限制了其广泛应用。本文提出ViewCrafter,一种从单张或稀疏图像出发,利用视频扩散模型先验,合成通用场景高保真新视角的新方法。该方法借助视频扩散模型的强大生成能力与基于点表示提供的粗略3D线索,生成高质量视频帧并实现精确相机姿态控制。为进一步扩大新视角生成范围,我们设计了迭代视角合成策略与相机轨迹规划算法,逐步扩展3D线索和覆盖区域。通过ViewCrafter,可高效优化3D-GS表示以支持实时沉浸式渲染,并实现场景级文本到3D生成,促进更具想象力的内容创作。在多个数据集上的实验表明,该方法具备强大的泛化能力与优越的生成性能。
原文摘要 · Abstract (English)
Despite recent advancements in neural 3D reconstruction, the dependence on dense multi-view captures restricts their broader applicability. In this work, we propose \textbf{ViewCrafter}, a novel method for synthesizing high-fidelity novel views of generic scenes from single or sparse images with the prior of video diffusion model. Our method takes advantage of the powerful generation capabilities of video diffusion model and the coarse 3D clues offered by point-based representation to generate high-quality video frames with precise camera pose control. To further enlarge the generation range of novel views, we tailored an iterative view synthesis strategy together with a camera trajectory planning algorithm to progressively extend the 3D clues and the areas covered by the novel views. With ViewCrafter, we can facilitate various applications, such as immersive experiences with real-time rendering by efficiently optimizing a 3D-GS representation using the reconstructed 3D points and the generated novel views, and scene-level text-to-3D generation for more imaginative content creation. Extensive experiments on diverse datasets demonstrate the strong generalization capability and superior performance of our method in synthesizing high-fidelity and consistent novel views.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。