arXiv:2512.06802cs.CV2025-12被引 2

用最优传输技术提升视频生成效率与质量,4步完成生成且效果媲美百步模型。

VDOT: Efficient Unified Video Creation via Optimal Transport Distillation

  • 引入最优传输优化分布匹配,提升训练稳定性和生成效率。
  • 4步生成效果媲美100步基线模型,显著缩短推理时间。
  • 支持多任务统一生成,适合需要高效视频创作的应用场景。

生成模型的快速发展推动了图像与视频应用的进步。其中,视频生成在多种条件下的生成备受关注,但现有模型或仅针对特定条件,或因复杂推理导致生成时间过长,难以实用。为此,我们提出高效统一视频生成模型VDOT。通过分布匹配蒸馏(DMD)框架,不采用传统的KL最小化,而是引入新型计算最优传输(OT)技术,优化真实与虚假得分分布间的差异。OT距离天然具有几何约束,可缓解少步生成中可能发生的零强制或梯度坍塌问题,从而提升蒸馏过程的效率与稳定性。进一步,引入判别器使模型感知真实视频数据,提升生成质量。为支持统一视频生成模型训练,我们设计全自动视频数据标注与过滤流水线,适用于多任务生成;同时构建统一测试基准UVCBench以标准化评估。实验表明,我们的4步VDOT在性能上优于或匹配其他具有100次去噪步骤的基线模型。

原文摘要 · Abstract (English)

The rapid development of generative models has significantly advanced image and video applications. Among these, video creation, aimed at generating videos under various conditions, has gained substantial attention. However, existing video creation models either focus solely on a few specific conditions or suffer from excessively long generation times due to complex model inference, making them impractical for real-world applications. To mitigate these issues, we propose an efficient unified video creation model, named VDOT. Concretely, we model the training process with the distribution matching distillation (DMD) paradigm. Instead of using the Kullback-Leibler (KL) minimization, we additionally employ a novel computational optimal transport (OT) technique to optimize the discrepancy between the real and fake score distributions. The OT distance inherently imposes geometric constraints, mitigating potential zero-forcing or gradient collapse issues that may arise during KL-based distillation within the few-step generation scenario, and thus, enhances the efficiency and stability of the distillation process. Further, we integrate a discriminator to enable the model to perceive real video data, thereby enhancing the quality of generated videos. To support training unified video creation models, we propose a fully automated pipeline for video data annotation and filtering that accommodates multiple video creation tasks. Meanwhile, we curate a unified testing benchmark, UVCBench, to standardize evaluation. Experiments demonstrate that our 4-step VDOT outperforms or matches other baselines with 100 denoising steps.

视频生成扩散模型最优传输

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。