提升视频生成中3D相机控制的精度与质量,实现更自然的镜头运动。
AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers

- 从低频特性出发优化相机位姿训练策略,加速收敛并改善画质。
- 仅在部分网络层注入相机条件,参数减少4倍,视觉质量提升10%。
- 引入2万条静态镜头动态视频数据集,增强相机与场景运动区分能力。
近期大量工作将3D相机控制融入基础文生视频模型,但生成结果常存在相机控制不准、画质下降的问题。本文从第一性原理分析视频中的相机运动特性,发现其具有低频特征,据此优化训练与推理时的位姿条件调度,加速训练收敛并提升视觉与运动质量。通过探测无条件视频扩散模型的表征,发现其隐含执行相机位姿估计,且仅部分层包含相机信息,由此提出仅在子网络注入相机条件,避免干扰其他视频特征,实现4倍参数压缩、训练速度提升,并带来10%的视觉质量改进。此外,补充了包含2万条多样动态视频的精选数据集(静止相机拍摄),帮助模型区分相机运动与场景运动,显著提升生成视频的姿态控制动态表现。综合上述发现,提出先进的3D相机控制架构AC3D,成为当前具备相机控制能力的生成视频模型新基准。
原文摘要 · Abstract (English)
Numerous works have recently integrated 3D camera control into foundational text-to-video models, but the resulting camera control is often imprecise, and video generation quality suffers. In this work, we analyze camera motion from a first principles perspective, uncovering insights that enable precise 3D camera manipulation without compromising synthesis quality. First, we determine that motion induced by camera movements in videos is low-frequency in nature. This motivates us to adjust train and test pose conditioning schedules, accelerating training convergence while improving visual and motion quality. Then, by probing the representations of an unconditional video diffusion transformer, we observe that they implicitly perform camera pose estimation under the hood, and only a sub-portion of their layers contain the camera information. This suggested us to limit the injection of camera conditioning to a subset of the architecture to prevent interference with other video features, leading to a 4x reduction of training parameters, improved training speed, and 10% higher visual quality. Finally, we complement the typical dataset for camera control learning with a curated dataset of 20K diverse, dynamic videos with stationary cameras. This helps the model distinguish between camera and scene motion and improves the dynamics of generated pose-conditioned videos. We compound these findings to design the Advanced 3D Camera Control (AC3D) architecture, the new state-of-the-art model for generative video modeling with camera control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。