arXiv:2412.05633cs.CV2024-12

提出高效连续视频流模型,显著降低生成延迟和计算量。

Efficient Continuous Video Flow Model for Video Prediction

  • 采用新方法建模多步生成过程,减少采样步骤
  • 模型尺寸减至原三分之一,预测1.5K帧视频更高效
  • 在多个数据集上达到顶尖性能,适合长视频生成

多步预测模型如扩散模型和修正流模型已成为生成任务的领先方案,但其生成新帧时的采样延迟高于单步方法。这一延迟在视频预测任务中尤为突出,因典型60秒视频包含约1500帧。本文提出一种新型多步过程建模方法,有效缓解延迟瓶颈,促进此类方法在视频预测中的应用。该方法不仅减少预测下一帧所需的采样步数,还将模型规模缩减至原始大小的三分之一。我们在KTH、BAIR动作机器人、Human3.6M和UCF101等标准视频预测数据集上进行了评估,结果表明该方法在这些基准上实现了领先性能。

原文摘要 · Abstract (English)

Multi-step prediction models, such as diffusion and rectified flow models, have emerged as state-of-the-art solutions for generation tasks. However, these models exhibit higher latency in sampling new frames compared to single-step methods. This latency issue becomes a significant bottleneck when adapting such methods for video prediction tasks, given that a typical 60-second video comprises approximately 1.5K frames. In this paper, we propose a novel approach to modeling the multi-step process, aimed at alleviating latency constraints and facilitating the adaptation of such processes for video prediction tasks. Our approach not only reduces the number of sample steps required to predict the next frame but also minimizes computational demands by reducing the model size to one-third of the original size. We evaluate our method on standard video prediction datasets, including KTH, BAIR action robot, Human3.6M and UCF101, demonstrating its efficacy in achieving state-of-the-art performance on these benchmarks.

视频预测扩散模型高效生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。