arXiv:2512.07821cs.CVcs.AI2025-12被引 2

生成4D视频,让画面在时空上保持几何与运动一致

WorldReel: 4D Video Generation with Consistent Geometry and Motion Modeling

  • 用显式4D表示同时生成图像与场景结构、相机轨迹和光流
  • 在大幅形变和镜头移动下仍保持几何一致性,减少视图时间伪影
  • 融合合成与真实数据训练,兼顾几何精度与视觉真实感

近期视频生成模型虽具高度逼真度,但在三维空间上仍存在根本性不一致。本文提出WorldReel,一种原生时空一致的4D视频生成器。该模型联合生成RGB帧与4D场景表征,包括点云图、相机轨迹和稠密光流映射,实现时间上的几何与外观建模一致性。其显式4D表示强制维持单一持续存在的场景,即使在大范围非刚性形变和显著相机运动下仍保持稳定。训练中结合合成数据(提供精确4D监督)与真实视频(增强视觉多样性与真实感),使模型可泛化至真实场景并保持强几何保真度。大量实验表明,WorldReel在动态场景与移动相机条件下达到新基准,显著提升几何一致性、运动连贯性,并降低视图-时间伪影。我们相信,WorldReel推动了视频生成向4D一致世界建模迈进,支持智能体通过统一稳定的时空表征进行渲染、交互与推理。

原文摘要 · Abstract (English)

Recent video generators achieve striking photorealism, yet remain fundamentally inconsistent in 3D. We present WorldReel, a 4D video generator that is natively spatio-temporally consistent. WorldReel jointly produces RGB frames together with 4D scene representations, including pointmaps, camera trajectory, and dense flow mapping, enabling coherent geometry and appearance modeling over time. Our explicit 4D representation enforces a single underlying scene that persists across viewpoints and dynamic content, yielding videos that remain consistent even under large non-rigid motion and significant camera movement. We train WorldReel by carefully combining synthetic and real data: synthetic data providing precise 4D supervision (geometry, motion, and camera), while real videos contribute visual diversity and realism. This blend allows WorldReel to generalize to in-the-wild footage while preserving strong geometric fidelity. Extensive experiments demonstrate that WorldReel sets a new state-of-the-art for consistent video generation with dynamic scenes and moving cameras, improving metrics of geometric consistency, motion coherence, and reducing view-time artifacts over competing methods. We believe that WorldReel brings video generation closer to 4D-consistent world modeling, where agents can render, interact, and reason about scenes through a single and stable spatiotemporal representation.

4D生成视频生成几何一致性运动建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。