用感知信息提升长视频生成的稳定性和质量
WorldWeaver: Generating Long-Horizon Video Worlds via Rich Perception
- 联合建模图像与感知条件,增强时间一致性
- 利用深度信息构建记忆库,减少运动漂移
- 分段噪声调度降低计算开销,适合长视频生成
生成式视频建模已取得显著进展,但在长序列中保持结构和时间一致性仍是挑战。现有方法主要依赖RGB信号,导致物体结构和运动随时间累积误差。为此,我们提出WorldWeaver,一种在统一长时序框架下联合建模RGB帧与感知条件的鲁棒框架。训练框架具备三大优势:首先,通过统一表示联合预测感知条件与颜色信息,显著提升时间一致性和运动动态;其次,利用深度线索(比RGB更抗漂移)构建记忆库,保留更清晰的上下文信息,提升长视频生成质量;第三,采用分段噪声调度训练预测组,进一步缓解漂移并降低计算成本。在基于扩散模型和修正流模型的广泛实验中,WorldWeaver有效减少了时间漂移,提升了生成视频保真度。
原文摘要 · Abstract (English)
Generative video modeling has made significant strides, yet ensuring structural and temporal consistency over long sequences remains a challenge. Current methods predominantly rely on RGB signals, leading to accumulated errors in object structure and motion over extended durations. To address these issues, we introduce WorldWeaver, a robust framework for long video generation that jointly models RGB frames and perceptual conditions within a unified long-horizon modeling scheme. Our training framework offers three key advantages. First, by jointly predicting perceptual conditions and color information from a unified representation, it significantly enhances temporal consistency and motion dynamics. Second, by leveraging depth cues, which we observe to be more resistant to drift than RGB, we construct a memory bank that preserves clearer contextual information, improving quality in long-horizon video generation. Third, we employ segmented noise scheduling for training prediction groups, which further mitigates drift and reduces computational cost. Extensive experiments on both diffusion- and rectified flow-based models demonstrate the effectiveness of WorldWeaver in reducing temporal drift and improving the fidelity of generated videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。