解决长视频生成中依赖建模与误差累积难题。
Pack and Force Your Memory: Long-form and Consistent Video Generation
- 提出MemoryPack,用文本和图像联合建模长短时依赖。
- 引入Direct Forcing,单步逼近训练-推理对齐,减少误差传播。
- 支持分钟级一致性生成,适合长视频应用开发。
长视频生成面临双重挑战:模型需捕捉长程依赖,同时避免自回归解码中的误差累积。为此,我们提出两项贡献。首先,针对动态上下文建模,提出可学习的上下文检索机制MemoryPack,利用文本和图像信息作为全局引导,联合建模短时与长时依赖,实现分钟级时间一致性。该设计随视频长度线性扩展,保持计算效率。其次,为缓解误差累积,提出Direct Forcing,一种高效的单步近似策略,提升训练-推理对齐性,从而抑制推理过程中的误差传播。MemoryPack与Direct Forcing协同作用,显著提升长视频生成的上下文一致性和可靠性,推动自回归视频模型的实际应用。
原文摘要 · Abstract (English)
Long-form video generation presents a dual challenge: models must capture long-range dependencies while preventing the error accumulation inherent in autoregressive decoding. To address these challenges, we make two contributions. First, for dynamic context modeling, we propose MemoryPack, a learnable context-retrieval mechanism that leverages both textual and image information as global guidance to jointly model short- and long-term dependencies, achieving minute-level temporal consistency. This design scales gracefully with video length, preserves computational efficiency, and maintains linear complexity. Second, to mitigate error accumulation, we introduce Direct Forcing, an efficient single-step approximating strategy that improves training-inference alignment and thereby curtails error propagation during inference. Together, MemoryPack and Direct Forcing substantially enhance the context consistency and reliability of long-form video generation, advancing the practical usability of autoregressive video models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。