无需相机位姿,用固定虚拟视角生成4D视频内容。
See4D: Pose-Free 4D Generation via Auto-Regressive Video Inpainting
- 用固定虚拟摄像机替代动态轨迹,解耦相机与场景建模。
- 通过视图条件化修复实现跨视角一致的4D生成,支持任意视角扩展。
- 适用于野外视频,无需3D标注,适合真实场景4D内容生成。
沉浸式应用需要从随意视频中合成时空4D内容,而无需昂贵的3D监督。现有视频到4D的方法通常依赖手动标注的相机位姿,成本高且对野外视频不稳定。近期的“先扭曲后修复”方法通过沿新相机轨迹扭曲输入帧并使用修复模型填补缺失区域,从而从多视角呈现4D场景。然而,这种轨迹到轨迹的范式常将相机运动与场景动态混淆,增加建模与推理复杂度。我们提出See4D,一种无位姿、轨迹到相机的框架,以固定虚拟摄像机阵列替代显式轨迹预测,实现相机控制与场景建模分离。训练一个视图条件化的视频修复模型,通过去噪真实合成的扭曲图像来学习鲁棒几何先验,并在不同虚拟视角间修复遮挡或缺失区域,消除对显式3D标注的需求。基于此修复核心,设计了时空自回归推理流程,沿虚拟摄像机样条遍历并扩展视频,采用重叠窗口机制,在每步保持有限复杂度。我们在跨视角视频生成和稀疏重建基准上验证了See4D,定量与定性评估均显示其优于依赖位姿或轨迹的基线方法,推动了从普通视频中实现实用4D世界建模。
原文摘要 · Abstract (English)
Immersive applications call for synthesizing spatiotemporal 4D content from casual videos without costly 3D supervision. Existing video-to-4D methods typically rely on manually annotated camera poses, which are labor-intensive and brittle for in-the-wild footage. Recent warp-then-inpaint approaches mitigate the need for pose labels by warping input frames along a novel camera trajectory and using an inpainting model to fill missing regions, thereby depicting the 4D scene from diverse viewpoints. However, this trajectory-to-trajectory formulation often entangles camera motion with scene dynamics and complicates both modeling and inference. We introduce See4D, a pose-free, trajectory-to-camera framework that replaces explicit trajectory prediction with rendering to a bank of fixed virtual cameras, thereby separating camera control from scene modeling. A view-conditional video inpainting model is trained to learn a robust geometry prior by denoising realistically synthesized warped images and to inpaint occluded or missing regions across virtual viewpoints, eliminating the need for explicit 3D annotations. Building on this inpainting core, we design a spatiotemporal autoregressive inference pipeline that traverses virtual-camera splines and extends videos with overlapping windows, enabling coherent generation at bounded per-step complexity. We validate See4D on cross-view video generation and sparse reconstruction benchmarks. Across quantitative metrics and qualitative assessments, our method achieves superior generalization and improved performance relative to pose- or trajectory-conditioned baselines, advancing practical 4D world modeling from casual videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。