用4D重建模型隐式几何知识,让单目视频重渲染更准更稳。
LaVR: Scene Latent Conditioned Generative Video Trajectory Re-Rendering using Large 4D Reconstruction Models
- 用大模型潜空间隐含结构代替显式深度图,避免误差累积
- 联合潜变量与相机位姿,生成视角变化时无漂移的视频
- 适合做高质量视频重渲染的研究者和工业应用
给定单目视频,视频重渲染的目标是生成新视角下的场景视图。现有方法面临两大挑战:几何无关模型缺乏空间感知,导致视角变化时产生漂移和形变;几何相关模型依赖估计深度与显式重建,易受深度误差和标定错误影响。本文提出利用大型4D重建模型潜空间中嵌入的隐式几何知识来条件化视频生成过程。这些潜变量在连续空间中捕捉场景结构,无需显式重建,提供灵活表示,使预训练扩散先验能更有效正则化误差。通过联合条件化潜变量与源相机位姿,我们在视频重渲染任务上达到当前最优性能。项目主页:https://lavr-4d-scene-rerender.github.io/。
原文摘要 · Abstract (English)
Given a monocular video, the goal of video re-rendering is to generate views of the scene from a novel camera trajectory. Existing methods face two distinct challenges. Geometrically unconditioned models lack spatial awareness, leading to drift and deformation under viewpoint changes. On the other hand, geometrically-conditioned models depend on estimated depth and explicit reconstruction, making them susceptible to depth inaccuracies and calibration errors. We propose to address these challenges by using the implicit geometric knowledge embedded in the latent space of a large 4D reconstruction model to condition the video generation process. These latents capture scene structure in a continuous space without explicit reconstruction. Therefore, they provide a flexible representation that allows the pretrained diffusion prior to regularize errors more effectively. By jointly conditioning on these latents and source camera poses, we demonstrate that our model achieves state-of-the-art results on the video re-rendering task. Project webpage is https://lavr-4d-scene-rerender.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。