arXiv:2510.19755cs.LGcs.AI2025-10综述被引 17

提出缓存机制,让扩散模型生成更快,无需重训练。

A Survey on Cache Methods in Diffusion Models: Toward Efficient Multi-Modal Generation

  • 通过复用扩散过程中的重复计算,实现无训练加速。
  • 从静态缓存发展到动态预测,适应更多任务场景。
  • 适合需要实时生成的多模态应用开发者参考。

扩散模型因其卓越的生成质量和可控性已成为现代生成式AI的核心。然而,其固有的多步迭代和复杂骨干网络带来了巨大的计算开销和生成延迟,成为实时应用的主要瓶颈。尽管现有加速技术有所进展,但仍面临适用性有限、训练成本高或质量下降等问题。在此背景下,扩散缓存(Diffusion Caching)提供了一种无需训练、与架构无关且高效的推理范式。其核心机制是识别并重用扩散过程中内在的计算冗余,通过特征级跨步骤复用和层间调度,减少计算量而不修改模型参数。本文系统梳理了扩散缓存的理论基础与演进历程,提出了统一分类与分析框架。通过对比代表性方法,揭示其从静态复用向动态预测发展的趋势,提升了在多样化任务中的灵活性,并可与采样优化、模型蒸馏等技术融合,为未来多模态与交互式应用构建统一高效的推理体系铺平道路。我们认为该范式将成为实现实时高效生成式AI的关键驱动力,为高效生成智能的理论与实践注入新活力。

原文摘要 · Abstract (English)

Diffusion Models have become a cornerstone of modern generative AI for their exceptional generation quality and controllability. However, their inherent \textit{multi-step iterations} and \textit{complex backbone networks} lead to prohibitive computational overhead and generation latency, forming a major bottleneck for real-time applications. Although existing acceleration techniques have made progress, they still face challenges such as limited applicability, high training costs, or quality degradation. Against this backdrop, \textbf{Diffusion Caching} offers a promising training-free, architecture-agnostic, and efficient inference paradigm. Its core mechanism identifies and reuses intrinsic computational redundancies in the diffusion process. By enabling feature-level cross-step reuse and inter-layer scheduling, it reduces computation without modifying model parameters. This paper systematically reviews the theoretical foundations and evolution of Diffusion Caching and proposes a unified framework for its classification and analysis. Through comparative analysis of representative methods, we show that Diffusion Caching evolves from \textit{static reuse} to \textit{dynamic prediction}. This trend enhances caching flexibility across diverse tasks and enables integration with other acceleration techniques such as sampling optimization and model distillation, paving the way for a unified, efficient inference framework for future multimodal and interactive applications. We argue that this paradigm will become a key enabler of real-time and efficient generative AI, injecting new vitality into both theory and practice of \textit{Efficient Generative Intelligence}.

扩散模型缓存加速生成效率多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。