用记忆库实现长期一致的虚拟世界模拟,支持跨视角重建。
WorldMem: Long-term Consistent World Simulation with Memory

- 引入记忆单元存储场景帧与状态信息,构建长期记忆库。
- 在视角或时间跨度较大时仍能准确还原过往场景。
- 适合需要动态世界建模的自动驾驶与机器人研究者。
世界模拟因能建模虚拟环境并预测行为后果而日益受到关注。然而,有限的时间上下文窗口常导致难以维持长期一致性,尤其在保持三维空间一致性方面表现不佳。本文提出WorldMem框架,通过包含记忆帧与状态(如位姿和时间戳)的记忆库,利用记忆注意力机制从历史记忆中有效提取相关信息。该方法可在显著视点或时间间隔下准确重构先前观测到的场景。同时,通过将时间戳纳入状态信息,框架不仅能建模静态世界,还能捕捉其随时间演化的动态过程,支持模拟世界中的感知与交互。在虚拟与真实场景中的大量实验验证了该方法的有效性。
原文摘要 · Abstract (English)
World simulation has gained increasing popularity due to its ability to model virtual environments and predict the consequences of actions. However, the limited temporal context window often leads to failures in maintaining long-term consistency, particularly in preserving 3D spatial consistency. In this work, we present WorldMem, a framework that enhances scene generation with a memory bank consisting of memory units that store memory frames and states (e.g., poses and timestamps). By employing a memory attention mechanism that effectively extracts relevant information from these memory frames based on their states, our method is capable of accurately reconstructing previously observed scenes, even under significant viewpoint or temporal gaps. Furthermore, by incorporating timestamps into the states, our framework not only models a static world but also captures its dynamic evolution over time, enabling both perception and interaction within the simulated world. Extensive experiments in both virtual and real scenarios validate the effectiveness of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。