arXiv:2508.21019cs.CV2025-08AAAI被引 6

提出新方法实现大模型视频生成单步加速,质量不降反升。

Phased One-Step Adversarial Equilibrium for Video Diffusion Models

  • 分两阶段训练:先对齐真实与生成视频分布,再通过自对抗机制优化
  • 在VBench-I2V上平均质量提升5.8%,大模型延迟降低100倍
  • 适合需要快速生成高质量视频的场景,如长视频创作或实时应用

视频扩散生成面临采样效率瓶颈,尤其在大规模模型和长时序任务中。现有加速方法多源自图像领域,缺乏对大规模视频模型的单步蒸馏能力,且难以泛化到条件下游任务。为此,我们提出视频分阶段对抗均衡(V-PAE)蒸馏框架,实现从大规模视频模型中高保真单步生成。该方法包含两阶段:(i) 稳定性预热阶段,对齐真实与生成视频分布,提升后续单步对抗蒸馏的稳定性;(ii) 统一对抗均衡阶段,复用生成器参数作为判别器主干,实现高斯噪声空间中的共进化对抗平衡。针对条件任务,有效缓解图像到视频生成中的语义退化与条件帧坍缩问题,保持视频-图像主体一致性。在VBench-I2V上的全面实验表明,V-PAE相较现有方法平均质量得分提升5.8%,涵盖语义对齐、时序连贯性与帧质量;同时将大模型(如Wan2.1-I2V-14B)扩散延迟降低100倍,仍保持竞争力性能。

原文摘要 · Abstract (English)

Video diffusion generation suffers from critical sampling efficiency bottlenecks, particularly for large-scale models and long contexts. Existing video acceleration methods, adapted from image-based techniques, lack a single-step distillation ability for large-scale video models and task generalization for conditional downstream tasks. To bridge this gap, we propose the Video Phased Adversarial Equilibrium (V-PAE), a distillation framework that enables high-quality, single-step video generation from large-scale video models. Our approach employs a two-phase process. (i) Stability priming is a warm-up process to align the distributions of real and generated videos. It improves the stability of single-step adversarial distillation in the following process. (ii) Unified adversarial equilibrium is a flexible self-adversarial process that reuses generator parameters for the discriminator backbone. It achieves a co-evolutionary adversarial equilibrium in the Gaussian noise space. For the conditional tasks, we primarily preserve video-image subject consistency, which is caused by semantic degradation and conditional frame collapse during the distillation training in image-to-video (I2V) generation. Comprehensive experiments on VBench-I2V demonstrate that V-PAE outperforms existing acceleration methods by an average of 5.8% in the overall quality score, including semantic alignment, temporal coherence, and frame quality. In addition, our approach reduces the diffusion latency of the large-scale video model (e.g., Wan2.1-I2V-14B) by 100 times, while preserving competitive performance.

视频生成扩散模型加速

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。