无需训练即可控制视频生成运动,通过对象变形和光流实现精细操控
QWERTY: Training-Free Motion Control via Query-Warped Video Diffusion Transformers

- 用查询变形和光流操纵3D注意力,实现训练自由的运动控制
- 在预训练图像到视频扩散变换器上达到最佳控制效果,接近微调性能
- 适合需要快速精确控制视频运动的研究者和开发者
视频扩散变换器(DiTs)能生成高保真且时间连贯的视频,但运动控制仍依赖文本提示,需反复调整。虽可通过添加边界框或轨迹等空间提示进行微调以实现显式控制,但代价高昂且可能损害生成能力。现有方法在基于U-Net的模型中探索了无训练的空间提示控制,但在DiTs中尚未深入。本文提出QWERTY,一种无需训练的框架,利用用户定义的对象变形与光流,在预训练图像到视频的DiTs中实现灵活运动控制。通过扭曲查询的3D不变语义子空间,操纵其3D全注意力机制。我们发现,经变形查询预测的噪声自然引导扩散过程朝向目标运动,并进一步利用该噪声作为自指导优化潜在表示,提升了控制稳定性和视觉质量。实验表明,QWERTY在最新图像到视频DiT上,相比现有无训练方法实现了最有效的运动控制,性能接近微调基线。
原文摘要 · Abstract (English)
Video diffusion transformers (DiTs) generate high-fidelity and temporally coherent videos, yet motion control remains implicit, primarily relying on text prompts. As a result, achieving desired motion often requires extensive prompt engineering and repeated resampling. While fine-tuning models with additional spatial prompts (e.g., bounding boxes or point trajectories) enables explicit control, it demands substantial data curation and computation, and may compromise the generative capabilities of pretrained models. Consequently, training-free motion control using such spatial prompts has been explored in U-Net-based video diffusion models, but remains largely unexplored for DiTs. We introduce QWERTY, a training-free framework that enables flexible motion control in pretrained image-to-video DiTs via user-defined object warping and optical flow. We carefully manipulate the 3D full attention of DiTs by warping the frame-invariant semantic subspace of queries. We find that the noise predicted by the query-warped DiT naturally guides the diffusion trajectory toward the desired motion, and further show that leveraging this noise as self-guidance for latent optimization improves control stability and visual quality. Experiments show that QWERTY achieves the most effective motion control among existing training-free approaches on a recent image-to-video DiT, with performance comparable to fine-tuning-based methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。