用隐空间流匹配实现视频插值与外推,支持任意帧率生成。
Video Latent Flow Matching: Optimal Polynomial Projections for Video Interpolation and Extrapolation
- 基于隐空间流建模,利用预训练图像模型生成时序视频帧。
- 理论证明投影误差有界且对时间尺度鲁棒,支持任意帧率插值外推。
- 适用于需要高精度视频生成的场景,如影视特效、动画制作。
本文提出一种高效的视频建模方法——视频隐空间流匹配(VLFM)。不同于以往随机采样隐空间块的方法,本方法依托强大的预训练图像生成模型,构建受文本引导的隐空间块时序流动,可解码为随时间变化的视频帧。我们假设视频的多帧在某一隐空间中对时间可微分,基于此引入HiPPO框架,逼近多项式最优投影以生成概率路径。该方法具备理论保障的有界通用近似误差和时间尺度鲁棒性。此外,VLFM可实现任意帧率下的视频插值与外推。我们在多个文生视频数据集上进行实验,验证了方法的有效性。
原文摘要 · Abstract (English)
This paper considers an efficient video modeling process called Video Latent Flow Matching (VLFM). Unlike prior works, which randomly sampled latent patches for video generation, our method relies on current strong pre-trained image generation models, modeling a certain caption-guided flow of latent patches that can be decoded to time-dependent video frames. We first speculate multiple images of a video are differentiable with respect to time in some latent space. Based on this conjecture, we introduce the HiPPO framework to approximate the optimal projection for polynomials to generate the probability path. Our approach gains the theoretical benefits of the bounded universal approximation error and timescale robustness. Moreover, VLFM processes the interpolation and extrapolation abilities for video generation with arbitrary frame rates. We conduct experiments on several text-to-video datasets to showcase the effectiveness of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。