arXiv:2506.05284cs.CV2025-06NeurIPS被引 124

用3D空间记忆增强视频世界模型的长期一致性

Video World Models with Long-term Spatial Memory

论文配图:Video World Models with Long-term Spatial Memory
图 1 · 摘自论文原文
  • 引入几何基准的长期空间记忆机制,存储环境结构信息
  • 在自回归视频生成中实现更长上下文和更强场景一致性
  • 适合需要长期视觉一致性的虚拟场景生成任务

新兴的世界模型通过自回归方式根据动作(如相机移动)和文本提示等控制信号生成视频帧。由于时间上下文窗口有限,这些模型在重访场景时经常出现严重遗忘,导致环境不一致。受人类记忆机制启发,我们提出一种新框架,通过基于几何的长期空间记忆来提升视频世界模型的长期一致性。该框架包含存储与检索长期空间记忆的机制,并构建了定制数据集以训练和评估具备显式3D记忆能力的世界模型。实验表明,相比基线方法,本方法在生成质量、一致性及上下文长度方面均有显著提升,为实现长期一致的世界生成铺平道路。

原文摘要 · Abstract (English)

Emerging world models autoregressively generate video frames in response to actions, such as camera movements and text prompts, among other control signals. Due to limited temporal context window sizes, these models often struggle to maintain scene consistency during revisits, leading to severe forgetting of previously generated environments. Inspired by the mechanisms of human memory, we introduce a novel framework to enhancing long-term consistency of video world models through a geometry-grounded long-term spatial memory. Our framework includes mechanisms to store and retrieve information from the long-term spatial memory and we curate custom datasets to train and evaluate world models with explicitly stored 3D memory mechanisms. Our evaluations show improved quality, consistency, and context length compared to relevant baselines, paving the way towards long-term consistent world generation.

视频生成世界模型空间记忆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。