提升视频扩散模型的相机运动控制精度,解决大模型生成中镜头动作不准问题。
Boosting Camera Motion Control for Video Diffusion Transformers
- 提出基于无分类器引导的相机运动引导机制,显著增强镜头控制能力。
- 在DiT模型上实现相机运动精度提升超400%,突破传统架构限制。
- 设计稀疏相机控制流程,简化长视频镜头轨迹设定,通用性强。
扩散模型在视频生成质量上取得显著进展,但对相机位姿的精细控制仍具挑战。尽管基于U-Net的模型在相机控制方面表现良好,基于Transformer的扩散模型(DiT)——大规模视频生成的主流架构——却存在严重的相机运动精度退化问题。本文深入分析该现象根源,发现相机控制性能主要取决于条件输入方式而非相机位姿表示形式。为解决这一问题,我们提出相机运动引导(CMG),基于无分类器引导,在DiT架构上使相机控制精度提升超过400%。此外,我们设计了一种稀疏相机控制流水线,极大简化了长视频中相机位姿的指定过程。所提方法可通用适配U-Net与DiT模型,广泛提升视频生成任务中的相机控制能力。
原文摘要 · Abstract (English)
Recent advancements in diffusion models have significantly enhanced the quality of video generation. However, fine-grained control over camera pose remains a challenge. While U-Net-based models have shown promising results for camera control, transformer-based diffusion models (DiT)-the preferred architecture for large-scale video generation - suffer from severe degradation in camera motion accuracy. In this paper, we investigate the underlying causes of this issue and propose solutions tailored to DiT architectures. Our study reveals that camera control performance depends heavily on the choice of conditioning methods rather than camera pose representations that is commonly believed. To address the persistent motion degradation in DiT, we introduce Camera Motion Guidance (CMG), based on classifier-free guidance, which boosts camera control by over 400%. Additionally, we present a sparse camera control pipeline, significantly simplifying the process of specifying camera poses for long videos. Our method universally applies to both U-Net and DiT models, offering improved camera control for video generation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。