arXiv:2603.03482cs.CVcs.AI2026-03中稿 · ICML被引 6

让世界模型具备持久3D状态,实现更真实连贯的交互生成。

Beyond Pixel Histories: World Models with Persistent 3D State

论文配图:Beyond Pixel Histories: World Models with Persistent 3D State
图 1 · 摘自论文原文
  • 构建可演化潜空间3D场景,包含环境、相机与渲染器。
  • 相比旧方法,空间记忆与3D一致性显著提升,长时稳定性增强。
  • 支持从单图生成多样3D环境,可在3D空间直接编辑与控制。

交互式世界模型通过响应用户操作持续生成视频,实现开放生成能力。然而现有模型通常缺乏环境的3D表示,导致3D一致性需从数据中隐式学习,且空间记忆受限于短时上下文窗口,造成用户体验不真实,并阻碍下游任务如智能体训练。为此,我们提出PERSIST,一种新范式的世界模型,模拟潜空间3D场景(环境、相机、渲染器)的演化过程。该机制支持生成具有持久空间记忆和一致几何结构的新帧。定量指标与定性用户研究均显示,相比现有方法,其在空间记忆、3D一致性及长时序稳定性方面均有显著提升,可生成连贯演化的3D世界。我们进一步展示新能力:仅凭单张图像即可合成多样化3D环境;支持在3D空间中直接进行环境编辑与指定,实现精细、基于几何的生成控制。

原文摘要 · Abstract (English)

Interactive world models continually generate video by responding to a user's actions, enabling open-ended generation capabilities. However, existing models typically lack a 3D representation of the environment, meaning 3D consistency must be implicitly learned from data, and spatial memory is restricted to limited temporal context windows. This results in an unrealistic user experience and presents significant obstacles to downstream tasks such as training agents. To address this, we present PERSIST, a new paradigm of world model which simulates the evolution of a latent 3D scene: environment, camera, and renderer. This allows us to synthesise new frames with persistent spatial memory and consistent geometry. Both quantitative metrics and a qualitative user study show substantial improvements in spatial memory, 3D consistency, and long-horizon stability over existing methods, enabling coherent, evolving 3D worlds. We further demonstrate novel capabilities, including synthesising diverse 3D environments from a single image, as well as enabling fine-grained, geometry-aware control over generated experiences by supporting environment editing and specification directly in 3D space. Project page: https://francelico.github.io/persist.github.io

世界模型3D生成持久记忆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。