arXiv:2601.16296cs.CVcs.AI2026-01被引 1

让视频编辑保持连贯,多轮修改不跑偏。

Memory-V2V: Memory-Augmented Video-to-Video Diffusion for Consistent Multi-Turn Editing

  • 用外部记忆存储之前修改内容,作为后续生成的约束
  • 多轮编辑后仍保持画面一致性,视觉质量不下降
  • 适合需要反复修改长视频的创作者使用

视频到视频扩散模型在单次编辑中表现优异,但实际编辑流程通常是迭代进行的。现有模型将每轮编辑独立处理,导致先前生成的区域出现漂移或被覆盖。我们识别出这一问题为多轮编辑中的跨轮一致性缺失。为此提出Memory-V2V框架,将先前编辑结果作为结构化约束,通过外部记忆存储输出、检索相关修改,并利用感知相关性的标记化与自适应压缩进行融合。该方法实现可扩展的条件控制,计算开销无线性增长。我们在迭代式视频新视角合成与文本引导的长视频编辑任务上验证了效果:Memory-V2V显著提升跨轮一致性,同时保持高质量视觉表现,优于强基线模型且仅带来适度额外开销。

原文摘要 · Abstract (English)

Video-to-video diffusion models achieve impressive single-turn editing performance, but practical editing workflows are inherently iterative. When edits are applied sequentially, existing models treat each turn independently, often causing previously generated regions to drift or be overwritten. We identify this failure mode as the problem of cross-turn consistency in multi-turn video editing. We introduce Memory-V2V, a memory-augmented framework that treats prior edits as structured constraints for subsequent generations. Memory-V2V maintains an external memory of previous outputs, retrieves task-relevant edits, and integrates them through relevance-aware tokenization and adaptive compression. These technical ingredients enable scalable conditioning without linear growth in computation. We demonstrate Memory-V2V on iterative video novel view synthesis and text-guided long video editing. Memory-V2V substantially enhances cross-turn consistency while maintaining visual quality, outperforming strong baselines with modest overhead.

视频编辑扩散模型一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。