Flowception通过交替插入与去噪帧,实现高效长视频生成。
Flowception: Temporally Expansive Flow Matching for Video Generation
- 交替进行帧插入与连续去噪,构建非自回归生成路径。
- 训练计算量减少3倍,FVD与VBench指标优于基线方法。
- 支持图像转视频、视频插值等多任务,可学习视频长度。
我们提出 Flowception,一种新型的非自回归、可变长度视频生成框架。该方法学习一条概率路径,将离散帧插入与连续帧去噪交替进行。相比自回归方法,采样过程中帧插入机制有效缓解误差累积,作为高效的长程上下文压缩手段。相比全序列流模型,本方法训练时计算量降低三倍,更适配局部注意力结构,并能联合学习视频长度与内容。定量实验表明,其在 FVD 与 VBench 指标上均优于自回归与全序列基线方法,定性结果也进一步验证了性能优势。通过学习序列中帧的插入与去噪,Flowception可无缝集成图像到视频生成、视频插值等任务。
原文摘要 · Abstract (English)
We present Flowception, a novel non-autoregressive and variable-length video generation framework. Flowception learns a probability path that interleaves discrete frame insertions with continuous frame denoising. Compared to autoregressive methods, Flowception alleviates error accumulation/drift as the frame insertion mechanism during sampling serves as an efficient compression mechanism to handle long-term context. Compared to full-sequence flows, our method reduces FLOPs for training three-fold, while also being more amenable to local attention variants, and allowing to learn the length of videos jointly with their content. Quantitative experimental results show improved FVD and VBench metrics over autoregressive and full-sequence baselines, which is further validated with qualitative results. Finally, by learning to insert and denoise frames in a sequence, Flowception seamlessly integrates different tasks such as image-to-video generation and video interpolation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。