用少量视角生成连贯自然的视频,无需训练。
Novel View Synthesis from A Few Glimpses via Test-Time Natural Video Completion
- 利用预训练视频扩散模型,在测试时补全缺失视角。
- 在极端稀疏输入下,重建质量显著优于现有方法。
- 适合无场景训练数据的3D重建与视频生成任务。
仅凭场景的少数视角,能否想象出摄像机滑过时的完整动态画面?我们提出将稀疏输入的新视角合成问题重新定义为测试时的自然视频补全任务,借助预训练视频扩散模型的强大先验来生成合理的中间视角。提出的零样本、生成引导框架,通过不确定性感知机制生成新视角下的伪图像,并以3D高斯点阵(3D-GS)作为几何基础,增强对观测不足区域的重建。通过迭代反馈,3D结构与2D生成相互优化,实现高质量、空间一致的渲染结果。该方法无需任何场景特定训练或微调,在LLFF、DTU、DL3DV和MipNeRF-360数据集上,极端稀疏条件下显著超越主流3D-GS基线。
原文摘要 · Abstract (English)
Given just a few glimpses of a scene, can you imagine the movie playing out as the camera glides through it? That's the lens we take on \emph{sparse-input novel view synthesis}, not only as filling spatial gaps between widely spaced views, but also as \emph{completing a natural video} unfolding through space. We recast the task as \emph{test-time natural video completion}, using powerful priors from \emph{pretrained video diffusion models} to hallucinate plausible in-between views. Our \emph{zero-shot, generation-guided} framework produces pseudo views at novel camera poses, modulated by an \emph{uncertainty-aware mechanism} for spatial coherence. These synthesized frames densify supervision for \emph{3D Gaussian Splatting} (3D-GS) for scene reconstruction, especially in under-observed regions. An iterative feedback loop lets 3D geometry and 2D view synthesis inform each other, improving both the scene reconstruction and the generated views. The result is coherent, high-fidelity renderings from sparse inputs \emph{without any scene-specific training or fine-tuning}. On LLFF, DTU, DL3DV, and MipNeRF-360, our method significantly outperforms strong 3D-GS baselines under extreme sparsity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。