arXiv:2512.25075cs.CVcs.AI2025-12被引 7

让视频生成同时控制视角和动作,实现任意时空探索。

SpaceTimePilot: Generative Rendering of Dynamic Scenes Across Space and Time

  • 分离时空维度,用时间嵌入显式控制运动序列。
  • 自研时序扭曲训练法,使模型学会精确时序操控。
  • 首个全空间时序覆盖数据集,适合影视与虚拟现实应用。

我们提出SpaceTimePilot,一种可分离时空的视频扩散模型,能基于单目视频独立调整摄像机视角与场景运动序列,实现跨时空的连续自由重渲染。为实现此目标,我们在扩散过程中引入有效的时间嵌入机制,实现对输出视频运动序列的显式控制。由于缺乏包含连续时序变化的配对动态视频数据集,我们提出一种简单但有效的时序扭曲训练策略,利用现有多视角数据集模拟时序差异,从而有效监督模型学习时序控制并实现鲁棒的时空解耦。为进一步提升双控精度,我们引入两个新组件:改进的相机条件机制,支持从首帧改变相机;以及首个合成的时空全覆盖渲染数据集CamxTime,提供场景内完全自由的时空视频轨迹。在时序扭曲训练方案与CamxTime数据集上联合训练后,模型展现出更精确的时序控制能力。我们在真实世界与合成数据上评估该模型,验证了清晰的时空解耦效果,并显著优于先前方法。

原文摘要 · Abstract (English)

We present SpaceTimePilot, a video diffusion model that disentangles space and time for controllable generative rendering. Given a monocular video, SpaceTimePilot can independently alter the camera viewpoint and the motion sequence within the generative process, re-rendering the scene for continuous and arbitrary exploration across space and time. To achieve this, we introduce an effective animation time-embedding mechanism in the diffusion process, allowing explicit control of the output video's motion sequence with respect to that of the source video. As no datasets provide paired videos of the same dynamic scene with continuous temporal variations, we propose a simple yet effective temporal-warping training scheme that repurposes existing multi-view datasets to mimic temporal differences. This strategy effectively supervises the model to learn temporal control and achieve robust space-time disentanglement. To further enhance the precision of dual control, we introduce two additional components: an improved camera-conditioning mechanism that allows altering the camera from the first frame, and CamxTime, the first synthetic space-and-time full-coverage rendering dataset that provides fully free space-time video trajectories within a scene. Joint training on the temporal-warping scheme and the CamxTime dataset yields more precise temporal control. We evaluate SpaceTimePilot on both real-world and synthetic data, demonstrating clear space-time disentanglement and strong results compared to prior work. Project page: https://zheninghuang.github.io/Space-Time-Pilot/ Code: https://github.com/ZheningHuang/spacetimepilot

视频生成扩散模型时空控制动态渲染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。