通过预测下一帧速率,高效生成长视频。
TempoMaster: Efficient Long Video Generation via Next-Frame-Rate Prediction
- 分阶段生成:先低帧率粗略构建,再逐步提升帧率细化
- 支持长时序连贯性,生成速度比现有方法快2倍以上
- 适合需要高质量长视频的场景,如影视创作、动画生成
我们提出TempoMaster,将长视频生成建模为下一帧速率预测问题。首先生成低帧率片段作为整个视频序列的粗略蓝图,随后逐步提升帧率以细化视觉细节和运动连续性。生成过程中,每个帧率层级内使用双向注意力,跨帧率则进行自回归,从而在保证长时序一致性的同时实现高效并行合成。大量实验表明,TempoMaster在长视频生成上达到新SOTA,视觉与时间质量均显著领先。
原文摘要 · Abstract (English)
We present TempoMaster, a novel framework that formulates long video generation as next-frame-rate prediction. Specifically, we first generate a low-frame-rate clip that serves as a coarse blueprint of the entire video sequence, and then progressively increase the frame rate to refine visual details and motion continuity. During generation, TempoMaster employs bidirectional attention within each frame-rate level while performing autoregression across frame rates, thus achieving long-range temporal coherence while enabling efficient and parallel synthesis. Extensive experiments demonstrate that TempoMaster establishes a new state-of-the-art in long video generation, excelling in both visual and temporal quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。