首次实现实时交互式4D视频生成,可同步控制相机与物体运动。
4DStreamCtrl: Interactive Video Generation with Online 4D Control

- 用3D点轨迹统一表示相机、物体运动和深度,支持联合控制。
- 单次前向传播完成相机、物体控制、深度编辑与动作迁移。
- 支持任意长度视频流生成,20帧/秒,适合交互式应用。
生成式视频模型已能合成近乎真实的画面。其作为交互工具的潜力取决于对物体与摄像机随时间运动的细粒度控制,但现有方法仅覆盖部分能力:摄像机参数方法可调整视角却无法移动物体,2D轨迹方法仅在图像平面操作而忽略深度与遮挡,近期3D方法虽引入几何信息,但仅支持离线固定长度生成。尤其缺乏兼具3D一致性、相机与物体联合控制且支持实时流式生成的方法。本文提出将相机运动、物体轨迹与深度统一为单一3D点轨迹表示,使单一模型在一次前向传播中完成相机与物体控制、深度编辑及动作迁移。为规模化学习该接口,我们从真实视频中挖掘3D运动监督信号,构建OpenVidHD-Motion3D数据集,并设计轻量级几何运动头,嵌入预训练视频扩散模型。由于该编码器具有时间可分离性,我们进一步将其蒸馏为因果流式学生模型,在内存独立于视频长度的条件下,以四步去噪实现任意长度视频生成。该统一设计在运动控制精度上超越此前仅控制相机、2D或离线3D的方法,涵盖它们各自单独处理的模态。4DStreamCtrl在单块高端GPU上以20 FPS运行480p视频,数百帧内保持时序连贯性,首次实现交互式4D可控流式生成。更广泛地,通过显式3D几何建模与高效因果推理,为具备闭环时空控制能力的交互式世界模型(如可控模拟器、具身智能体的实时视觉想象)指明方向。
原文摘要 · Abstract (English)
Generative video models now synthesize footage nearly indistinguishable from reality. Their promise as interactive tools hinges on fine-grained control of how objects and the camera move over time, yet each existing approach captures only part of this: camera-parameter methods steer the viewpoint but cannot move objects, 2D-trajectory methods act in the image plane and ignore depth and occlusion, and recent 3D methods add geometry but run only offline at a fixed length. In particular, none combines 3D-consistent control of both camera and objects with real-time, streaming generation. Here we show that camera motion, object trajectories, and depth can be unified into a single 3D point-track representation, from which one model performs joint camera and object control, depth editing, and motion transfer in a single forward pass. To learn this interface at scale, we mine in-the-wild video for 3D motion supervision, yielding OpenVidHD-Motion3D, and encode it with a lightweight Geometric Motion Head that plugs into a pretrained video diffusion model. Because this encoder is temporally separable, we distill the model into a causal streaming student that generates arbitrarily long video in four denoising steps at memory independent of length. This unified design surpasses prior camera-only, 2D, and offline-3D methods in motion-control precision while covering modalities they address only in isolation. 4DStreamCtrl runs at 20 FPS on a single high-end GPU for 480p video and stays temporally coherent over hundreds of frames, enabling, to our knowledge, interactive 4D-controllable streaming generation for the first time. More broadly, grounding generation in explicit 3D geometry with efficient causal inference points toward interactive world models with closed-loop spatiotemporal control, from controllable simulators to real-time visual imagination for embodied agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。