通过智能代理选择加速视频生成,效率提升超2倍且画质几乎无损。
GalaxyDiT: Efficient Video Generation with Guidance Alignment and Adaptive Proxy in Diffusion Transformers
- 利用相关性分析选定最优代理指标,实现计算资源高效复用。
- 在Wan2.1模型上达1.87×~2.37×加速,画质下降不足1%。
- 无需重新训练,适合追求高效生成的视频应用开发者。
扩散模型已彻底改变视频生成领域,成为创意内容生成与物理模拟的核心工具。基于Transformer的架构(DiTs)和无分类器引导(CFG)是其成功的关键,能有效遵循提示并生成高质量视频。然而,这类模型计算开销巨大:每次生成需数十次迭代,且CFG使计算量翻倍,限制了其在下游应用中的普及。本文提出GalaxyDiT,一种无需训练的加速方法,结合引导对齐与自适应代理选择,系统性优化计算复用。通过秩相关分析,我们为不同模型家族与参数规模确定了最优代理指标,确保高效重用。在Wan2.1-1.3B和Wan2.1-14B模型上分别实现1.87×和2.37×加速,VBench-2.0基准测试中画质仅下降0.97%和0.72%。在高加速比下,仍保持优于基线模型5~10 dB的峰值信噪比(PSNR),显著超越现有最先进方法。
原文摘要 · Abstract (English)
Diffusion models have revolutionized video generation, becoming essential tools in creative content generation and physical simulation. Transformer-based architectures (DiTs) and classifier-free guidance (CFG) are two cornerstones of this success, enabling strong prompt adherence and realistic video quality. Despite their versatility and superior performance, these models require intensive computation. Each video generation requires dozens of iterative steps, and CFG doubles the required compute. This inefficiency hinders broader adoption in downstream applications. We introduce GalaxyDiT, a training-free method to accelerate video generation with guidance alignment and systematic proxy selection for reuse metrics. Through rank-order correlation analysis, our technique identifies the optimal proxy for each video model, across model families and parameter scales, thereby ensuring optimal computational reuse. We achieve $1.87\times$ and $2.37\times$ speedup on Wan2.1-1.3B and Wan2.1-14B with only 0.97% and 0.72% drops on the VBench-2.0 benchmark. At high speedup rates, our approach maintains superior fidelity to the base model, exceeding prior state-of-the-art approaches by 5 to 10 dB in peak signal-to-noise ratio (PSNR).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。