arXiv:2608.29910cs.CV2026-09

让虚拟世界实时互动更稳定,支持长时间一致的场景与相机控制。

Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory

论文配图:Matrix-Game 3.5: Enhancing Real-Time Streaming Interactive World Models with Patch Memory
图 1 · 摘自论文原文
  • 用无额外参数的内存框架融合3D贴图检索与相机投影,保持几何一致性。
  • 分离静态场景与动态物体建模,确保长期生成中对象身份不变。
  • 通过两阶段蒸馏实现分钟级实时交互生成,适合游戏与元宇宙应用。

交互式世界模型将视频生成从离线片段合成扩展到持久模拟交互式虚拟世界,适用于游戏、机器人、具身智能体和扩展现实(XR)。然而,实现稳定的长时程交互生成仍具挑战,需同时维持场景几何、动态一致性及相机控制,并支持实时自回归生成。基于Matrix-Game 3.0,我们提出Matrix-Game 3.5,通过三项关键改进推动实时交互世界生成向几何感知与长时程一致方向发展:第一,提出统一的几何感知记忆框架,其贴片内存(patch-memory)与平铺投影位置编码(tiled-PRoPE)不引入额外可学习参数,结合显式3D贴片检索与投影相机条件,实现几何一致的相机控制与忠实的长时程场景回溯;第二,引入静态-动态解耦的世界表征,分别建模静态场景几何与动态主体,保障长期生成中的几何一致性和主体身份;第三,开发两阶段渐进式实时蒸馏框架,通过感知流匹配(Perceptual Flow Matching)与课程式自滚动去噪模型(Self-Rollout DMD),将双向扩散模型转化为几步因果生成器,支持分钟级实时交互生成。大量实验表明,在涵盖Unreal仿真环境、开放世界游戏与互联网视频的统一训练数据集上,MatrixGame 3.5在长时程场景回溯、精确相机控制、主体一致性、提示驱动世界生成及稳定实时开放世界交互方面均表现优异。

原文摘要 · Abstract (English)

Interactive world models extend video generation from offline clip synthesis toward persistent simulation of interactive virtual worlds, enabling applications in games, robotics, embodied agents, and XR. Achieving stable long-horizon interactive generation, however, remains challenging, as the model must simultaneously preserve scene geometry, dynamic consistency, and camera control while supporting real-time autoregressive generation. Building upon Matrix-Game 3.0, we present Matrix-Game 3.5, as shown in Figure 1, which advances real-time interactive world generation toward geometry-aware and long-horizon consistent simulation through three key improvements. First, we propose a unified geometry-aware memory framework, whose patch-memory and tiled-PRoPE components introduce no additional learnable parameters, combining explicit 3D patch retrieval with projective camera conditioning to enable geometry-consistent camera control and faithful long-horizon scene recall. Second, we introduce a static-dynamic disentangled world representation that separately models static scene geometry and dynamic subjects, preserving both geometric consistency and subject identity throughout long-horizon generation. Third, we develop a two-stage progressive real-time distillation framework that converts a bidirectional diffusion model into a few-step causal generator through Perceptual Flow Matching and curriculum based Self-Rollout DMD, enabling minute-long real-time interactive generation. Extensive experiments demonstrate that, with a unified training corpus spanning Unreal simulation environments, open-world games, and internet videos, MatrixGame 3.5 achieves strong performance in long-horizon scene recall, precise camera control, subject consistency, prompt-driven world generation, and stable real-time open-world interaction.

世界模型实时生成几何一致交互系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。