arXiv:2505.19151cs.GRcs.AI2025-05被引 6

用大模型抓结构、小模型精细节,视频生成快三倍还不丢质量。

SRDiffusion: Accelerate Video Diffusion Inference via Sketching-Rendering Cooperation

  • 大模型负责高噪声阶段的语义和运动保持,小模型处理低噪声阶段的细节优化。
  • 在VBench评测中对Wan模型提速超3倍,质量几乎无损;对CogVideoX提速2倍。
  • 适合需要高效生成高清长视频的研究者与开发者,可与现有加速方法并用。

基于扩散变换器(DiT)架构的模型如Sora、CogVideoX和Wan,在文本到视频、图像到视频及视频编辑任务中取得了显著进展。然而,基于扩散的视频生成仍计算开销巨大,尤其在高分辨率、长时长视频生成时。以往工作通过跳过计算加速推理,但常导致质量严重下降。本文提出SRDiffusion框架,通过大模型与小模型协同降低推理成本:大模型在高噪声阶段确保语义与运动保真度(草图阶段),小模型在低噪声阶段优化视觉细节(渲染阶段)。实验表明,该方法优于现有方案,在VBench评测中对Wan实现超过3倍速度提升且质量几乎无损,对CogVideoX实现2倍提速。该方法为可扩展视频生成提供了一种与现有加速策略正交的新方向。

原文摘要 · Abstract (English)

Leveraging the diffusion transformer (DiT) architecture, models like Sora, CogVideoX and Wan have achieved remarkable progress in text-to-video, image-to-video, and video editing tasks. Despite these advances, diffusion-based video generation remains computationally intensive, especially for high-resolution, long-duration videos. Prior work accelerates its inference by skipping computation, usually at the cost of severe quality degradation. In this paper, we propose SRDiffusion, a novel framework that leverages collaboration between large and small models to reduce inference cost. The large model handles high-noise steps to ensure semantic and motion fidelity (Sketching), while the smaller model refines visual details in low-noise steps (Rendering). Experimental results demonstrate that our method outperforms existing approaches, over 3$\times$ speedup for Wan with nearly no quality loss for VBench, and 2$\times$ speedup for CogVideoX. Our method is introduced as a new direction orthogonal to existing acceleration strategies, offering a practical solution for scalable video generation.

视频生成扩散模型加速推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。