arXiv:2604.18215cs.CV2026-04被引 3

分离记忆与生成,让长时视频更连贯且能探索新场景。

Memorize When Needed: Decoupled Memory Control for Spatially Consistent Long-Horizon Video Generation

论文配图:Memorize When Needed: Decoupled Memory Control for Spatially Consistent Long-Horizon Video Generation
图 1 · 摘自论文原文
  • 用独立轻量记忆分支学习空间一致性,不干扰生成过程。
  • 跨帧注意力机制确保每帧只参考最相关的历史信息。
  • 相机感知门控机制让记忆只在必要时介入,适合新场景生成。

空间一致的长时视频生成旨在沿预设摄像机轨迹保持时空一致性。现有方法多将记忆建模与生成过程耦合,导致场景回溯时内容不一致,探索新区域时生成能力下降,即使使用大量标注数据训练也如此。为此,我们提出解耦框架,将记忆条件与生成分离。该方法显著降低训练成本,同时提升空间一致性并保留对新场景的生成能力。具体地,采用轻量级独立记忆分支,从历史观测中学习精确的空间一致性;引入混合记忆表示,融合生成帧中的时空互补线索,并通过逐帧交叉注意力机制,确保每帧仅基于最相关的过去信息进行条件化,再注入生成模型以保证空间一致性。生成新场景时,设计相机感知门控机制,仅在存在有意义历史参考时启用记忆条件作用。相比现有方法,本方法高度数据高效,实验表明其在视觉质量与空间一致性上均达到当前最优水平。

原文摘要 · Abstract (English)

Spatially consistent long-horizon video generation aims to maintain temporal and spatial consistency along predefined camera trajectories. Existing methods mostly entangle memory modeling with video generation, leading to inconsistent content during scene revisits and diminished generative capacity when exploring novel regions, even trained on extensive annotated data. To address these limitations, we propose a decoupled framework that separates memory conditioning from generation. Our approach significantly reduces training costs while simultaneously enhancing spatial consistency and preserving the generative capacity for novel scene exploration. Specifically, we employ a lightweight, independent memory branch to learn precise spatial consistency from historical observation. We first introduce a hybrid memory representation to capture complementary temporal and spatial cues from generated frames, then leverage a per-frame cross-attention mechanism to ensure each frame is conditioned exclusively on the most spatially relevant historical information, which is injected into the generative model to ensure spatial consistency. When generating new scenes, a camera-aware gating mechanism is proposed to mediate the interaction between memory and generation modules, enabling memory conditioning only when meaningful historical references exist. Compared with the existing method, our method is highly data-efficient, yet the experiments demonstrate that our approach achieves state-of-the-art performance in terms of both visual quality and spatial consistency.

视频生成空间一致性记忆解耦长时序列

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。