arXiv:2507.08422cs.CVeess.IV2025-07被引 9

无需训练的图像生成加速方法,可实现7倍提速且几乎不损失质量。

Training-free Mixed-Resolution Latent Upsampling for Spatially Accelerated Diffusion Transformers

  • 通过自适应地仅在边缘区域提前上采样,避免图像失真。
  • 在FLUX-1.dev上实现最高7.0倍加速,Stable Diffusion 3上达3.0倍。
  • 兼容已有时间维度加速技术,总提速可达15.9倍,适合部署优化。

扩散变压器(DiTs)在高保真生成中表现出卓越的可扩展性,但其计算开销对实际部署构成巨大挑战。现有加速方法主要利用时间维度,而空间加速仍缺乏探索。本文研究通过潜在空间上采样实现DiTs的空间加速。发现直接上采样会引入伪影,主要源于高频边缘区域的混叠及噪声-时间步差异导致的不匹配。基于此,提出无需训练的空间加速框架——区域自适应潜在上采样(RALU),通过仅在易出错边缘区域提前上采样,并实现不同潜在分辨率间的噪声-时间步匹配,有效缓解伪影。RALU在保持高质量的前提下,使FLUX-1.dev实现最高7.0×加速,Stable Diffusion 3达3.0×加速,且与现有时间加速方法和时间步蒸馏模型兼容,组合提速最高达15.9×。

原文摘要 · Abstract (English)

Diffusion transformers (DiTs) offer excellent scalability for high-fidelity generation, but their computational overhead poses a great challenge for practical deployment. Existing acceleration methods primarily exploit the temporal dimension, whereas spatial acceleration remains underexplored. In this work, we investigate spatial acceleration for DiTs via latent upsampling. We found that naïve latent upsampling for spatial acceleration introduces artifacts, primarily due to aliasing in high-frequency edge regions and mismatching from noise-timestep discrepancies. Then, based on these findings and analyses, we propose a training-free spatial acceleration framework, dubbed Region-Adaptive Latent Upsampling (RALU), to mitigate those artifacts while achieving spatial acceleration of DiTs by our mixed-resolution latent upsampling. RALU achieves artifact-free, efficient acceleration with early upsampling only on artifact-prone edge regions and noise-timestep matching for different latent resolutions, leading to up to 7.0$\times$ speedup on FLUX-1.dev and 3.0$\times$ on Stable Diffusion 3 with negligible quality degradation. Furthermore, our RALU is complementarily applicable to existing temporal acceleration methods and timestep-distilled models, leading to up to 15.9$\times$ speedup.

扩散模型图像生成加速推理潜在空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。