arXiv:2506.07177cs.CVcs.AI2025-06被引 9

无需训练即可实现视频生成的帧级精细控制,支持多种输入信号。

Frame Guidance: Training-Free Guidance for Frame-Level Control in Video Diffusion Models

  • 通过简单潜空间处理降低内存占用,实现无训练帧级控制。
  • 在不微调模型的前提下,对多类输入信号生成高质量可控视频。
  • 适用于关键帧、风格迁移、循环等任务,兼容任意视频扩散模型。

扩散模型的发展显著提升了视频质量,也推动了细粒度可控生成的关注。然而,现有方法多依赖对大规模视频模型进行微调以适配特定任务,随着模型规模持续增大,这一方式日益不可行。本文提出 Frame Guidance,一种基于帧级信号(如关键帧、风格参考图、草图或深度图)的无训练可控视频生成方法。为实现实用的无训练引导,我们设计了一种简单的潜空间处理方法,大幅降低内存开销,并提出一种面向全局一致性的新型潜空间优化策略。该方法可在无需训练的情况下,有效控制多种任务,包括关键帧引导、风格化和循环生成,且兼容任意视频扩散模型。实验表明,Frame Guidance 能够为广泛的任务和输入信号生成高质量可控视频。

原文摘要 · Abstract (English)

Advancements in diffusion models have significantly improved video quality, directing attention to fine-grained controllability. However, many existing methods depend on fine-tuning large-scale video models for specific tasks, which becomes increasingly impractical as model sizes continue to grow. In this work, we present Frame Guidance, a training-free guidance for controllable video generation based on frame-level signals, such as keyframes, style reference images, sketches, or depth maps. For practical training-free guidance, we propose a simple latent processing method that dramatically reduces memory usage, and apply a novel latent optimization strategy designed for globally coherent video generation. Frame Guidance enables effective control across diverse tasks, including keyframe guidance, stylization, and looping, without any training, compatible with any video models. Experimental results show that Frame Guidance can produce high-quality controlled videos for a wide range of tasks and input signals.

视频生成扩散模型无训练帧级控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。