arXiv:2607.02798cs.CVcs.GR2026-07

通过3D对齐的噪声空间,统一控制视频中物体运动与镜头视角。

Track the Noise, Move the World:3D-Grounded Motion-Consistent Noise for Controllable Video Generation

论文配图:Track the Noise, Move the World:3D-Grounded Motion-Consistent Noise for Controllable Video Generation
图 1 · 摘自论文原文
  • 构建共享3D对齐的噪声空间,结合点追踪与球形噪声实现运动一致
  • 在Wan 2.1模型上实现最先进画质与运动可控性,无需修改主模型结构
  • 仅用轻量LoRA微调,适合希望精确控制视频动态的开发者

当前图像与文本到视频的扩散模型通过迭代去噪生成高保真视频,但难以在单一生成过程中同时精确控制物体运动与相机视角。本文提出UniCaMo框架,通过直接构造扩散模型输入噪声,实现物体轨迹与摄像机视角的统一控制。该方法建立跨视频帧的3D对齐运动一致性噪声空间:利用稀疏3D点追踪将参考帧噪声沿预设物体轨迹扭曲,同时采用虚拟球形噪声表示为相机运动带来的新场景区域提供全局一致的噪声值。通过局部轨迹引导的噪声扭曲与全局球面采样相结合,UniCaMo在物体移动和视角变化下均保持几何与时间一致性。由于仅修改输入噪声,无需额外适配器、控制分支或模型架构调整。在大型预训练视频扩散模型(如Wan 2.1,14B)上仅用轻量LoRA微调,UniCaMo在标准可控视频生成基准上达到最优的视频质量与运动可控性表现。

原文摘要 · Abstract (English)

Modern image-and-text-to-video diffusion models can synthesize highly realistic videos by iteratively denoising an initial Gaussian noise tensor conditioned on reference image and text inputs. However, existing approaches still lack precise and unified controllability over both object motion and camera motion within a single generation process. We present UniCaMo, a unified framework that enables simultaneous control of object trajectories and camera viewpoints by directly constructing the input noise of the diffusion model. Specifically, UniCaMo builds a shared 3D-grounded motion-consistent noise space across latent video frames. Sparse 3D point tracks are used to warp the Gaussian noise of the reference frame along desired object trajectories, while a virtual spherical noise representation provides globally consistent noise values for newly revealed scene regions under camera motion. By combining local track-guided noise warping with global sphere-based noise sampling, UniCaMo maintains geometric and temporal consistency under both object movement and viewpoint changes. Because UniCaMo modifies only the input noise, it requires no auxiliary adapters, control branches, or architectural changes to the underlying video diffusion model. With lightweight LoRA fine-tuning on large pretrained video diffusion models, including Wan 2.1 (14B), UniCaMo achieves state-of-the-art results in both video quality and motion controllability on standard controllable video generation benchmarks.

视频生成3D控制扩散模型运动一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。