arXiv:2507.12646cs.CV2025-07NeurIPS被引 15

用单目视频实现动态场景新视角合成,三步法兼顾精度与效率。

Reconstruct, Inpaint, Test-Time Finetune: Dynamic Novel-view Synthesis from Monocular Videos

论文配图:Reconstruct, Inpaint, Test-Time Finetune: Dynamic Novel-view Synthesis from Monocular Videos
图 1 · 摘自论文原文
  • 先重建3D动态场景,再用2D扩散模型补全不可见区域。
  • 在多个数据集上超越现有方法,新视角图像质量显著提升。
  • 无需标注数据,可零样本适配新视频,适合真实场景应用。

我们研究从单目视频中进行动态场景的新视角合成。以往方法依赖昂贵的测试时优化4D表示,或在前馈训练下无法保持场景几何结构。我们的方法基于三个关键洞察:(1)输入与目标视图共可见像素可通过先重建动态3D场景并从新视角渲染重建结果来生成;(2)新视角中隐藏像素可由前馈2D视频扩散模型“修复”;(3)我们的视频修复扩散模型CogNVS可从2D视频自监督训练,从而在大量真实视频上进行训练。这使得CogNVS可通过测试时微调零样本应用于新测试视频。实验证明,CogNVS在单目视频动态场景新视角合成任务中优于几乎所有先前方法。

原文摘要 · Abstract (English)

We explore novel-view synthesis for dynamic scenes from monocular videos. Prior approaches rely on costly test-time optimization of 4D representations or do not preserve scene geometry when trained in a feed-forward manner. Our approach is based on three key insights: (1) covisible pixels (that are visible in both the input and target views) can be rendered by first reconstructing the dynamic 3D scene and rendering the reconstruction from the novel-views and (2) hidden pixels in novel views can be "inpainted" with feed-forward 2D video diffusion models. Notably, our video inpainting diffusion model (CogNVS) can be self-supervised from 2D videos, allowing us to train it on a large corpus of in-the-wild videos. This in turn allows for (3) CogNVS to be applied zero-shot to novel test videos via test-time finetuning. We empirically verify that CogNVS outperforms almost all prior art for novel-view synthesis of dynamic scenes from monocular videos.

新视角合成视频生成扩散模型单目视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。