arXiv:2503.04871cs.CVcs.LG2025-03被引 1

用轻量解码器加速图像视频生成,提速近20倍

Toward Lightweight and Fast Decoders for Diffusion Models in Image and Video Generation

  • 设计轻量Vision Transformer和Taming Transformer解码器替代原模型
  • 图像生成快15%,视频子模块快20倍,内存适度降低
  • 适合大规模生成场景,如批量生成10万张图像

我们研究通过引入轻量级解码器来减少稳定扩散模型在图像和视频生成中的推理时间和内存占用。传统潜在扩散流程依赖大型变分自编码器解码器,导致生成速度慢且消耗大量GPU内存。本文提出使用轻量级Vision Transformer和Taming Transformer架构进行定制训练的解码器。实验表明,在COCO2017上图像生成整体速度提升最高达15%,在子模块中视频解码速度提升高达20倍,UCF-101数据集上也获得额外加速。内存需求适度降低,尽管与默认解码器相比感知质量略有下降,但速度和可扩展性的提升对大规模推理场景(如生成10万张图像)至关重要。本工作还结合了高效视频生成的进展,如双掩码策略,体现了提升生成模型可扩展性与效率的总体趋势。

原文摘要 · Abstract (English)

We investigate methods to reduce inference time and memory footprint in stable diffusion models by introducing lightweight decoders for both image and video synthesis. Traditional latent diffusion pipelines rely on large Variational Autoencoder decoders that can slow down generation and consume considerable GPU memory. We propose custom-trained decoders using lightweight Vision Transformer and Taming Transformer architectures. Experiments show up to 15% overall speed-ups for image generation on COCO2017 and up to 20 times faster decoding in the sub-module, with additional gains on UCF-101 for video tasks. Memory requirements are moderately reduced, and while there is a small drop in perceptual quality compared to the default decoder, the improvements in speed and scalability are crucial for large-scale inference scenarios such as generating 100K images. Our work is further contextualized by advances in efficient video generation, including dual masking strategies, illustrating a broader effort to improve the scalability and efficiency of generative models.

扩散模型轻量化视频生成加速推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。