arXiv:2411.18664cs.CV2024-11CVPR被引 33

无需训练的视频生成引导方法,提升画质同时保持动作自然多样。

Spatiotemporal Skip Guidance for Enhanced Video Diffusion Sampling

  • 通过跳过时空层模拟弱模型,实现免训练引导
  • 在不损失运动多样性的前提下,显著提升生成视频质量
  • 适合追求高质量且动态自然视频生成的研究者

扩散模型已成为生成高质量图像、视频和3D内容的强大工具。尽管条件控制(CFG)等采样引导技术能提升质量,但会降低多样性与运动连贯性。自引导方法可缓解此问题,但需额外弱模型训练,限制其在大规模模型中的应用。本文提出时空跳过引导(STG),一种针对基于Transformer的视频扩散模型的无训练采样引导方法。STG通过自扰动隐式构建弱模型,无需外部模型或额外训练。通过选择性跳过时空层,生成与原模型对齐的退化版本,从而在不牺牲多样性与动态程度的前提下提升样本质量。主要贡献包括:(1) 提出高效高表现的视频扩散模型引导技术;(2) 通过层跳过模拟弱模型,消除对外部模型的需求;(3) 在不损害样本多样性与动态性的前提下实现质量增强,优于传统CFG。更多结果详见 https://junhahyung.github.io/STGuidance。

原文摘要 · Abstract (English)

Diffusion models have emerged as a powerful tool for generating high-quality images, videos, and 3D content. While sampling guidance techniques like CFG improve quality, they reduce diversity and motion. Autoguidance mitigates these issues but demands extra weak model training, limiting its practicality for large-scale models. In this work, we introduce Spatiotemporal Skip Guidance (STG), a simple training-free sampling guidance method for enhancing transformer-based video diffusion models. STG employs an implicit weak model via self-perturbation, avoiding the need for external models or additional training. By selectively skipping spatiotemporal layers, STG produces an aligned, degraded version of the original model to boost sample quality without compromising diversity or dynamic degree. Our contributions include: (1) introducing STG as an efficient, high-performing guidance technique for video diffusion models, (2) eliminating the need for auxiliary models by simulating a weak model through layer skipping, and (3) ensuring quality-enhanced guidance without compromising sample diversity or dynamics unlike CFG. For additional results, visit https://junhahyung.github.io/STGuidance.

视频生成扩散模型采样优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。