通过动态调节引导强度,实现文本到图像/视频生成的高效加速。
MAMBO-G: Magnitude-Aware Mitigation for Boosted Guidance
- 根据更新与预测的幅度比动态调整引导强度,稳定生成轨迹。
- 在SD3.5上提速3倍,在Lumina上提速4倍,14B参数视频模型提速2倍。
- 无需训练、即插即用,适合资源受限的大规模视频生成任务。
高保真文本到图像和文本到视频生成通常依赖无分类器引导(CFG),但获得最佳结果常需计算成本高昂的采样调度。本文提出MAMBO-G,一种无需训练的加速框架,通过动态优化引导幅度显著降低计算开销。我们发现标准CFG调度效率低下,早期步骤中施加了过大的更新,阻碍收敛速度。MAMBO-G基于更新与预测的幅度比调节引导尺度,有效稳定生成轨迹并实现快速收敛。该效率对资源密集型任务如视频生成尤为关键。方法为通用即插即用加速器,在Stable Diffusion v3.5(SD3.5)上实现最高3倍提速,在Lumina上达4倍;尤其在140亿参数的Wan2.1视频模型上提速2倍,同时保持视觉保真度。实现基于主流开源扩散框架,可无缝集成现有流程。
原文摘要 · Abstract (English)
High-fidelity text-to-image and text-to-video generation typically relies on Classifier-Free Guidance (CFG), but achieving optimal results often demands computationally expensive sampling schedules. In this work, we propose MAMBO-G, a training-free acceleration framework that significantly reduces computational cost by dynamically optimizing guidance magnitudes. We observe that standard CFG schedules are inefficient, applying disproportionately large updates in early steps that hinder convergence speed. MAMBO-G mitigates this by modulating the guidance scale based on the update-to-prediction magnitude ratio, effectively stabilizing the trajectory and enabling rapid convergence. This efficiency is particularly vital for resource-intensive tasks like video generation. Our method serves as a universal plug-and-play accelerator, achieving up to 3x speedup on Stable Diffusion v3.5 (SD3.5) and 4x on Lumina. Most notably, MAMBO-G accelerates the 14B-parameter Wan2.1 video model by 2x while preserving visual fidelity, offering a practical solution for efficient large-scale video synthesis. Our implementation follows a mainstream open-source diffusion framework and is plug-and-play with existing pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。