arXiv:2603.02943cs.CV2026-03被引 8

用帕德近似提升扩散模型采样速度,低步数下仍保持高质量生成。

TC-Padé: Trajectory-Consistent Padé Approximation for Diffusion Acceleration

  • 基于帕德近似建模特征演化,比泰勒方法更精准捕捉动态变化。
  • 在20-30步下实现1.72至2.88倍加速,各项质量指标稳定。
  • 适合需要快速生成且对质量要求高的图像视频应用。

尽管扩散模型已达到顶尖生成质量,但其迭代采样过程带来巨大计算负担。现有特征缓存技术在高步数(如50步)下有效,但在20-30步的实用低步数场景中表现受限。随着步距增大,泰勒类外推器易产生误差累积与轨迹漂移。同时,传统缓存策略忽视不同去噪阶段的动态差异。为此,本文提出轨迹一致的帕德近似(TC-Padé),基于有理函数建模特征演化,更准确捕捉渐近与过渡行为。为实现低步数下稳定、一致的采样,TC-Padé引入(1)自适应系数调制,利用历史缓存残差检测细微轨迹转变;(2)分阶段预测策略,适配早期、中期、晚期采样阶段的不同动态特性。在DiT-XL/2、FLUX.1-dev和Wan2.1上的大量实验表明,该方法显著优于现有缓存技术:例如在FLUX.1-dev上实现2.88倍加速,在Wan2.1上实现1.72倍加速,同时在FID、CLIP、Aesthetic及VBench-2.0等指标上保持高质量。

原文摘要 · Abstract (English)

Despite achieving state-of-the-art generation quality, diffusion models are hindered by the substantial computational burden of their iterative sampling process. While feature caching techniques achieve effective acceleration at higher step counts (e.g., 50 steps), they exhibit critical limitations in the practical low-step regime of 20-30 steps. As the interval between steps increases, polynomial-based extrapolators like TaylorSeer suffer from error accumulation and trajectory drift. Meanwhile, conventional caching strategies often overlook the distinct dynamical properties of different denoising phases. To address these challenges, we propose Trajectory-Consistent Padé approximation, a feature prediction framework grounded in Padé approximation. By modeling feature evolution through rational functions, our approach captures asymptotic and transitional behaviors more accurately than Taylor-based methods. To enable stable and trajectory-consistent sampling under reduced step counts, TC-Padé incorporates (1) adaptive coefficient modulation that leverages historical cached residuals to detect subtle trajectory transitions, and (2) step-aware prediction strategies tailored to the distinct dynamics of early, mid, and late sampling stages. Extensive experiments on DiT-XL/2, FLUX.1-dev, and Wan2.1 across both image and video generation demonstrate the effectiveness of TC-Padé. For instance, TC-Padé achieves 2.88x acceleration on FLUX.1-dev and 1.72x on Wan2.1 while maintaining high quality across FID, CLIP, Aesthetic, and VBench-2.0 metrics, substantially outperforming existing feature caching methods.

扩散模型加速生成帕德近似采样优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。