让视频生成同时控制时间与镜头运动,实现精准动态与视角调控。
BulletTime: Decoupled Control of Time and Camera Pose for Video Generation
- 分离场景动态与镜头运动,用4D位置编码和自适应归一化实现双控
- 在多样化时间模式和镜头轨迹下保持高质量生成,可控性超越现有方法
- 公开独立参数化的数据集,适合需要精细控制的视频生成研究者
新兴的视频扩散模型虽具备高视觉保真度,但将场景动态与镜头运动耦合,难以实现精确的空间与时间控制。本文提出一种4D可控视频扩散框架,显式解耦场景动态与相机姿态,支持对场景动态和摄像机视角的细粒度操控。该框架以连续世界时间序列和相机轨迹为条件输入,通过注意力层中的4D位置编码及特征调制的自适应归一化进行注入。为训练此模型,我们构建了一个独特的数据集,其中时间变化与相机运动独立参数化;该数据集将公开。实验表明,模型在多种时间模式和相机轨迹下均实现稳健的4D控制,同时保持高生成质量,并在可控性上优于先前工作。
原文摘要 · Abstract (English)
Emerging video diffusion models achieve high visual fidelity but fundamentally couple scene dynamics with camera motion, limiting their ability to provide precise spatial and temporal control. We introduce a 4D-controllable video diffusion framework that explicitly decouples scene dynamics from camera pose, enabling fine-grained manipulation of both scene dynamics and camera viewpoint. Our framework takes continuous world-time sequences and camera trajectories as conditioning inputs, injecting them into the video diffusion model through a 4D positional encoding in the attention layer and adaptive normalizations for feature modulation. To train this model, we curate a unique dataset in which temporal and camera variations are independently parameterized; this dataset will be made public. Experiments show that our model achieves robust real-world 4D control across diverse timing patterns and camera trajectories, while preserving high generation quality and outperforming prior work in controllability. See our website for codes and video results: https://19reborn.github.io/Bullet4D/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。