arXiv:2606.16449cs.CV2026-06

让视频编辑后保持长期一致,靠拆分视觉与结构记忆实现

PermaVid: Consistent Video Generation Across Edits via Disentangled Context Memory

论文配图:PermaVid: Consistent Video Generation Across Edits via Disentangled Context Memory
图 1 · 摘自论文原文
  • 分离外观与几何结构的双记忆库设计
  • 编辑后仍能保持语义和结构一致性,超越现有方法
  • 适合需要长期视频编辑的场景,如影视制作

视频编辑后的持续一致性要求模型在修改场景外观或布局后,仍能保持跨时间与视角的生成连贯性。然而,现有记忆机制在修改后难以维持长期一致性,因存储上下文可能过时或失效。为此,我们提出 PermaVid,一种基于多模态上下文记忆的新框架,将空间上下文解耦为语义外观与几何结构,并引入感知编辑的记忆更新与检索策略,使记忆演化与后续观测对齐。具体地,构建两个互补的记忆库:RGB 上下文记忆捕捉外观信息并隐式编码几何,深度上下文记忆则仅保留几何结构且与语义解耦。在此基础上,设计一种记忆引导的视频生成模型,在混合模态记忆上下文中进行多模态特征融合。实验表明,该方法在编辑后仍能保持强长期语义与结构一致性,显著优于当前最优方法。

原文摘要 · Abstract (English)

Consistent video generation under editing operations requires persistence: when edits modify scene appearance or layout, subsequent generations should remain coherent across time and viewpoints. However, existing memory designs struggle to maintain long-term consistency after such modifications, as stored contexts may become outdated or invalid. To address this, we propose PermaVid, a novel framework built upon a multi-modal context memory that disentangles spatial context into semantic appearance and geometric structure, together with an edit-aware memory update and retrieval strategy that keeps memory evolution aligned with subsequent observations. Specifically, we develop two complementary memory banks: an RGB context memory that captures appearance-aware observations while implicitly encoding geometry, and a depth context memory that preserves geometry-only structure disentangled from semantics. Building on this design, we introduce a memory-guided video generation model that performs multi-modal feature fusion under reference conditions drawn from mixed-modality memory contexts. Experiments demonstrate that our method maintains strong long-term semantic and structural consistency after edits, significantly outperforming state-of-the-art methods.

视频生成记忆机制一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。