提出可泛化动态视图合成新方法,支持高精度相机控制。
GRVS: a Generalizable and Recurrent Approach to Monocular Dynamic View Synthesis
- 采用循环结构实现输入与目标视频的异步无界映射。
- 在UCSD和Kubric-4D-dyn数据集上优于4种基于高斯溅射的方法。
- 支持六自由度精细相机控制,适用于复杂动态场景建模。
从单目动态场景视频中合成新视角仍具挑战性。依赖显式运动先验的场景特定方法在高度动态区域易失效,而基于扩散模型的方法虽视觉逼真但存在几何不一致问题,且计算开销大。受静态场景泛化模型启发,本文提出一种新模型,包含两个核心组件:(1) 循环结构实现输入与目标视频间的无界、异步映射;(2) 利用平面扫掠高效分离相机与场景运动,实现六自由度精细相机控制。模型在UCSD数据集及新构建的Kubric-4D-dyn数据集(包含更长、更高分辨率、更复杂动态序列)上训练评估。结果表明,该模型在静态与动态区域均优于四种基于高斯溅射的场景特定方法及两种扩散模型,在细粒度几何重建方面表现更优。
原文摘要 · Abstract (English)
Synthesizing novel views from monocular videos of dynamic scenes remains a challenging problem. Scene-specific methods that optimize 4D representations with explicit motion priors often break down in highly dynamic regions where multi-view information is hard to exploit. Diffusion-based approaches that integrate camera control into large pre-trained models can produce visually plausible videos but frequently suffer from geometric inconsistencies across both static and dynamic areas. Both families of methods also require substantial computational resources. Building on the success of generalizable models for static novel view synthesis, we adapt the framework to dynamic inputs and propose a new model with two key components: (1) a recurrent loop that enables unbounded and asynchronous mapping between input and target videos and (2) an efficient use of plane sweeps over dynamic inputs to disentangle camera and scene motion, and achieve fine-grained, six-degrees-of-freedom camera controls. We train and evaluate our model on the UCSD dataset and on Kubric-4D-dyn, a new monocular dynamic dataset featuring longer, higher resolution sequences with more complex scene dynamics than existing alternatives. Our model outperforms four Gaussian Splatting-based scene-specific approaches, as well as two diffusion-based approaches in reconstructing fine-grained geometric details across both static and dynamic regions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。