用历史画面当记忆,让视频生成更连贯。
MemCam: Memory-Augmented Camera Control for Consistent Video Generation
- 把已生成帧当作外部记忆,增强上下文信息
- 长视频大角度镜头下一致性显著提升
- 适合需要稳定画面的交互式视频创作
交互式视频生成在场景模拟和视频创作中潜力巨大,但现有方法在动态相机控制下的长时间生成中常因上下文信息不足导致场景不一致。为此,我们提出MemCam,一种基于记忆增强的交互式视频生成方法,将先前生成的帧作为外部记忆,并利用其作为上下文条件,实现可控相机视角下的高场景一致性生成。为支持更长、更相关的上下文,设计了上下文压缩模块,将记忆帧编码为紧凑表示,并采用基于共可见性的选择机制,动态检索最相关的历史帧,在降低计算开销的同时丰富上下文信息。在交互式视频生成任务上的实验表明,MemCam在场景一致性方面显著优于现有基线方法及开源最先进模型,尤其在包含大幅相机旋转的长视频场景中表现突出。
原文摘要 · Abstract (English)
Interactive video generation has significant potential for scene simulation and video creation. However, existing methods often struggle with maintaining scene consistency during long video generation under dynamic camera control due to limited contextual information. To address this challenge, we propose MemCam, a memory-augmented interactive video generation approach that treats previously generated frames as external memory and leverages them as contextual conditioning to achieve controllable camera viewpoints with high scene consistency. To enable longer and more relevant context, we design a context compression module that encodes memory frames into compact representations and employs co-visibility-based selection to dynamically retrieve the most relevant historical frames, thereby reducing computational overhead while enriching contextual information. Experiments on interactive video generation tasks show that MemCam significantly outperforms existing baseline methods as well as open-source state-of-the-art approaches in terms of scene consistency, particularly in long video scenarios with large camera rotations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。