用几何估计指导生成,实现单目视频的任意摄像机路径重编辑
Vid-CamEdit: Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry
- 先估时空一致的几何结构,再基于此生成新视角视频
- 在真实场景极端路径外推下仍能生成合理视频,优于现有方法
- 无需大量4D训练数据,适合处理野外实拍视频
我们提出Vid-CamEdit,一种新型视频摄像机轨迹编辑框架,可沿用户定义的相机路径重新合成单目视频。该任务因病态性及训练所需的多视角视频数据有限而具挑战性。传统重建方法难以应对极端轨迹变化,现有动态新视角合成生成模型无法处理真实场景视频。我们的方法分为两步:估计时序一致的几何结构,并基于此进行生成渲染。通过引入几何先验,生成模型聚焦于几何不确定性区域的细节合成。我们采用分解式微调框架,分别利用多视角图像和视频数据训练空间与时间组件,避免了对大规模4D训练数据的依赖。实验表明,本方法在真实世界视频上,尤其在极端外推场景下,生成视频的合理性显著优于基线方法。
原文摘要 · Abstract (English)
We introduce Vid-CamEdit, a novel framework for video camera trajectory editing, enabling the re-synthesis of monocular videos along user-defined camera paths. This task is challenging due to its ill-posed nature and the limited multi-view video data for training. Traditional reconstruction methods struggle with extreme trajectory changes, and existing generative models for dynamic novel view synthesis cannot handle in-the-wild videos. Our approach consists of two steps: estimating temporally consistent geometry, and generative rendering guided by this geometry. By integrating geometric priors, the generative model focuses on synthesizing realistic details where the estimated geometry is uncertain. We eliminate the need for extensive 4D training data through a factorized fine-tuning framework that separately trains spatial and temporal components using multi-view image and video data. Our method outperforms baselines in producing plausible videos from novel camera trajectories, especially in extreme extrapolation scenarios on real-world footage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。