arXiv:2605.25500cs.CV2026-05被引 2

从单视角视频生成完整动态4D场景,突破视点与内容限制。

Full-4D: Generating Full-Scope 4D Scenes from a Single-View Video

论文配图:Full-4D: Generating Full-Scope 4D Scenes from a Single-View Video
图 1 · 摘自论文原文
  • 构建多视角视频生成+优化重建框架,融合几何先验提升生成精度。
  • 在Real-MV-4D数据集上实现高质量同步多视角视频生成,支持全范围视点变化。
  • 采用流匹配蒸馏损失优化4D表示,适合需要真实动态场景还原的研究者。

从单视角视频生成4D场景本质上是病态问题:单一视角无法恢复完整且动态的全视野场景。现有方法通常局限于单目视频、简单3D效果或原视角附近的小范围视点偏移,难以实现真正的4D生成。同时,缺乏大规模同步多视角视频数据集也制约了该方向发展。本文提出一种新型单视角视频到4D的端到端框架,将全范围4D生成建模为多视角视频合成后接基于优化的4D重建。关键贡献包括:1)构建Real-MV-4D数据集,包含多样化真实环境中的同步多视角视频,提供4D监督信号;2)设计基于融合时间-视图注意力机制的多视角视频扩散模型,直接嵌入几何重投影先验与相机条件,使视图-时间交互严格对齐物理3D规律,生成稠密同步的T×V视频网格;3)不依赖非交互式且不一致的2D插值,而是将合成的多视角视频显式提升至4D表示(4DGS),并引入流匹配蒸馏损失,利用多视角先验增强新视角渲染质量。大量实验表明,本方法在视觉保真度与几何一致性上均优于现有方法,首次实现从单视角视频生成完整4D场景。

原文摘要 · Abstract (English)

Generating 4D scenes from a single-view video is inherently ill-posed: a single viewpoint lacks the information needed to recover a complete, dynamic scene with full coverage. Existing methods are typically limited to monocular videos, simple 3D effects, or only small viewpoint perturbations around the original viewpoint, falling short of true 4D generation. Meanwhile, the lack of large-scale datasets capturing full-scope 4D scenes with synchronized multi-view videos further hinders progress in this direction. We propose a novel single-view video-to-4D framework that casts full-scope 4D generation as a multi-view video synthesis followed by optimization-based 4D reconstruction from the generated views. To instantiate this formulation end-to-end, we make three key contributions. First, we introduce Real-MV-4D, a large-scale dataset of synchronized multi-view videos captured in diverse real-world environments to provide the 4D supervision. Second, we train a multi-view video diffusion model driven by a novel fused time(T)-view(V) attention mechanism that directly embeds geometric reprojection priors and explicit camera conditioning into its view-time interactions. Unlike basic feature fusion, this direct binding strictly aligns the generation process with physical 3D priors to produce a dense, synchronized T$\times $V video grid. Third, rather than relying on non-interactive and inconsistent 2D video interpolations, we lift the synthesized multi-view videos into an explicit 4D representation (i.e. 4DGS), regularized by a Flow Matching Distillation loss that exploits the multi-view prior to improve novel-view rendering. Extensive experiments demonstrate that our method outperforms existing approaches in both visual fidelity and geometric consistency, enabling full-scope 4D scene generation from single-view videos.

4D生成多视角视频扩散3D重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。