让视频世界模型学会自适应查询记忆,提升长期生成一致性。
MemLearner: Learning to Query Context memory for Video World Models

- 用查询标记学习动态选择上下文帧作为记忆
- 在遮挡和动态物体场景中显著提升画面一致性
- 无需额外训练模块,直接利用预训练模型能力
视频世界模型是基于用户操作和历史视频帧预测未来世界状态的交互式视频生成模型。其核心挑战在于缺乏有效记忆,导致长时间生成时场景不一致。以往方法采用规则化的上下文帧检索作为记忆,但在存在场景遮挡和动态物体的情况下泛化能力差。本文提出MemLearner,一种基于学习的自适应上下文查询方法,使用查询标记连接上下文与预测标记。通过利用视频生成模型自身进行上下文查询,MemLearner无需从零训练额外模块即可利用预训练视觉先验,并引入高效的训练与推理策略。我们构建了一个包含长视频、场景遮挡和动态物体并带有相机位姿标注的数据集,提出多数据集训练策略,融合带标注渲染视频与无标注真实视频。大量实验表明,MemLearner在场景一致性和记忆能力上显著优于先前视频世界模型,尤其在复杂遮挡与动态场景下表现更优。
原文摘要 · Abstract (English)
Video World Models are interactive video generation models that predict future world states based on user actions and history video frames. A critical challenge in video world models is the lack of memory, causing inconsistent generated scenes over extended durations. Previous methods explored rule-based context frame retrieval as memory, but they fail to generalize in scenarios with scene occlusions and dynamic objects. We propose MemLearner, a learning-based adaptive context query method using query tokens to bridge context and predicted tokens. By leveraging the video generation model itself for context querying, MemLearner exploits pre-trained visual priors without training additional modules from scratch, and incorporates efficient strategies for training and inference. We collect a dataset of long videos with scene occlusions and dynamic objects, paired with camera pose annotations, and propose a multi-dataset training strategy leveraging both annotated rendered and unannotated real-world videos. Extensive experiments demonstrate that MemLearner significantly outperforms prior video world models in terms of scene consistency and memory, particularly under challenging occlusion and dynamic scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。