解决长视频生成中物体持久性与记忆容量的矛盾问题。
Ring Forcing: Towards Precise Long-Term Memory for Autoregressive Video Diffusion

- 采用环形训练策略,强制模型从远期历史中检索信息。
- 在固定序列长度下实现分钟级有效历史跨度,提升长期一致性。
- 引入稀疏RoPE机制,灵活扩展记忆并保留预训练先验知识。
将视频生成扩展到长时间段时,现有模型存在严重的长期记忆缺陷。这一缺陷体现在两个关键方面:物体持久性(物体重新出现时能精确复现其外观)和记忆容量(处理超长上下文并利用遥远历史信息的能力)。稳健的长期记忆需要两者兼备:仅有持久性而缺乏上下文处理会限制时间范围,仅具备长上下文但无持久性则无法维持身份识别。为此,我们提出Ring Forcing,一种自回归视频扩散框架,旨在稳健构建并精确利用长期记忆。其环形结构训练策略强制从遥远历史中检索信息,有效调和严格历史遵循与生成多样性之间的权衡。为扩展记忆容量,我们引入压缩与时间步组合策略,在固定序列长度约束下,将有效历史跨度扩展至分钟级,并实现对整个历史的全面感知场。此外,我们提出稀疏RoPE机制,实现灵活可扩展的记忆适配,同时充分利用预训练先验。大量实验表明,Ring Forcing在分钟级连贯性和物体持久性上均显著优于现有最先进方法。
原文摘要 · Abstract (English)
Scaling video generation to long durations reveals a critical bottleneck: current models lack robust long-term memory. This deficiency can be studied along two critical aspects: object permanence, the ability to precisely reproduce the appearance of objects upon re-entry; and memory capacity, the ability to process ultra-long context and use information from distant history. Robust long-term memory requires both: object permanence without sufficient context handling limits the temporal scope, while long context length without permanence fails to maintain identity. To address this, we present Ring Forcing, an autoregressive video diffusion framework designed to robustly construct and precisely utilize long-term memory. Our ring-structured training strategy enforces retrieval from distant history, effectively reconciling the trade-off between strict historical adherence and generative diversity. To expand memory capacity, we introduce a compression and timestep composition strategy. Under fixed sequence length constraints, this method extends the effective historical span to minutes-long durations and achieves a comprehensive receptive field over the entire history. Furthermore, we present a sparse RoPE mechanism to enable flexible, scalable memory adaptation while fully exploiting pre-trained priors. Extensive experiments demonstrate that Ring Forcing achieves superior minutes-long coherence and object permanence, significantly outperforming state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。