arXiv:2601.18993cs.CVcs.AI2026-01International Conf…被引 4

无需训练,通过完整前景4D重建实现单目视频任意视角重播

FreeOrbit4D: Training-Free Arbitrary Camera Redirection for Monocular Videos via Foreground-Complete 4D Reconstruction

  • 分离前景与背景重建,用多视角扩散模型补全完整前景点云
  • 在大角度视角变化下仍保持几何一致性和时间连贯性,生成更真实视频
  • 适合需要高质量视角重播的视频编辑、4D数据生成等场景

相机重定向旨在从单目视频中沿用户指定的相机轨迹重放动态场景。然而,大角度重定向本质上是病态问题:单目视频仅捕捉动态3D场景的狭窄时空视图,对底层4D世界的信息极为有限。核心挑战在于从这种有限输入中恢复出完整且一致的表示,确保几何与运动的一致性。尽管近期基于扩散的方法在视觉生成质量上表现优异,但在远离原始轨迹的大角度视角变化下常因缺乏视觉锚点而出现严重几何模糊和时间不一致。本文提出FreeOrbit4D,一种无需训练的框架,通过恢复包含完整前景的4D代理作为生成的结构锚点来解决这一模糊性。该代理通过解耦前景与背景重建获得:将单目视频反投影至统一全局空间中的静态背景与部分前景点云,再利用以物体为中心的多视角扩散模型合成多视角图像,并在规范物体空间中重建完整的前景点云。通过密集像素同步的3D-3D对应关系将规范前景点云对齐到全局场景空间,并将完整的4D代理投影至目标相机视角,为条件视频扩散模型提供几何骨架。大量实验表明,FreeOrbit4D在具有挑战性的大角度轨迹下生成更忠实、时间更连贯的重播视频,且该代理还支持编辑传播和4D数据生成等应用。

原文摘要 · Abstract (English)

Camera redirection aims to replay a dynamic scene from a single monocular video under a user-specified camera trajectory. However, large-angle redirection is inherently ill-posed: a monocular video captures only a narrow spatio-temporal view of a dynamic 3D scene, providing severely limited observations of the underlying 4D world. The key challenge is therefore to recover a complete and coherent representation from this limited input, with consistent geometry and motion. While recent diffusion-based methods achieve impressive visual generation quality, they often break down under large-angle viewpoint changes far from the original trajectory, where missing visual grounding leads to severe geometric ambiguity and temporal inconsistency. We present FreeOrbit4D, an effective training-free framework that tackles this ambiguity by recovering a foreground-complete 4D proxy as structural grounding for video generation. We obtain this proxy by decoupling foreground and background reconstructions: we unproject the monocular video into a static background and partial foreground point clouds in a unified global space, then use an object-centric multi-view diffusion model to synthesize multi-view images and reconstruct complete foreground point clouds in canonical object space. By aligning the canonical foreground point cloud to the global scene space via dense pixel-synchronized 3D-3D correspondences and projecting the foreground-complete 4D proxy onto target camera viewpoints, we provide geometric scaffolds that guide a conditional video diffusion model. Extensive experiments show that FreeOrbit4D produces more faithful and temporally coherent redirected videos under challenging large-angle trajectories, and our proxy further enables applications such as edit propagation and 4D data generation. Project page: https://freeorbit4d.vision.ischool.illinois.edu/

视角重播4D重建扩散模型单目视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。