让轨迹可控视频生成从多步变少步,又快又准。
FlashMotion: Few-Step Controllable Video Generation with Trajectory Guidance
- 先训多步模型控轨迹,再蒸馏成少步版本加速
- 少步生成视频质量与轨迹精度均超越现有方法
- 适合需要快速生成精准运动视频的研究者
轨迹可控视频生成近年进展显著,但主流方法依赖多步去噪过程,导致时间冗余和计算开销大。现有视频蒸馏方法虽能将多步模型压缩为少步,但直接用于轨迹可控生成时,视频质量和轨迹准确性明显下降。为此,我们提出FlashMotion,一种面向少步轨迹可控视频生成的新训练框架:首先在多步生成器上训练轨迹适配器以实现精确轨迹控制;然后将生成器蒸馏为少步版本以加速生成;最后采用扩散与对抗联合优化的微调策略,使适配器与少步生成器对齐,生成高质量且轨迹准确的视频。我们还构建了FlashBench基准,评估长序列下多种前景物体的视频质量与轨迹一致性。在两种适配器架构上的实验表明,FlashMotion在视觉质量与轨迹一致性上均优于现有蒸馏方法及先前多步模型。
原文摘要 · Abstract (English)
Recent advances in trajectory-controllable video generation have achieved remarkable progress. Previous methods mainly use adapter-based architectures for precise motion control along predefined trajectories. However, all these methods rely on a multi-step denoising process, leading to substantial time redundancy and computational overhead. While existing video distillation methods successfully distill multi-step generators into few-step, directly applying these approaches to trajectory-controllable video generation results in noticeable degradation in both video quality and trajectory accuracy. To bridge this gap, we introduce FlashMotion, a novel training framework designed for few-step trajectory-controllable video generation. We first train a trajectory adapter on a multi-step video generator for precise trajectory control. Then, we distill the generator into a few-step version to accelerate video generation. Finally, we finetune the adapter using a hybrid strategy that combines diffusion and adversarial objectives, aligning it with the few-step generator to produce high-quality, trajectory-accurate videos. For evaluation, we introduce FlashBench, a benchmark for long-sequence trajectory-controllable video generation that measures both video quality and trajectory accuracy across varying numbers of foreground objects. Experiments on two adapter architectures show that FlashMotion surpasses existing video distillation methods and previous multi-step models in both visual quality and trajectory consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。