实现720p实时长视频生成,支持分钟级记忆一致性。
Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory
- 用合成数据+游戏采集+真实视频增强,构建大规模四元组数据集。
- 通过残差建模与相机感知记忆检索,实现分钟级时序一致性。
- 多段自回归蒸馏+量化剪枝,5B模型达40帧/秒实时生成。
随着交互式视频生成的发展,扩散模型在世界模型应用中展现出巨大潜力。然而,现有方法仍难以同时实现带记忆的长期时序一致性和高分辨率实时生成,限制了其在真实场景中的应用。为此,我们提出Matrix-Game 3.0,一种面向720p实时长视频生成的记忆增强型交互世界模型。在Matrix-Game 2.0基础上,我们在数据、模型和推理层面进行系统性改进:首先,构建升级版工业级无限数据引擎,融合基于Unreal Engine的合成数据、来自AAA游戏的大规模自动化采集及真实视频增强,规模化生成高质量视频-姿态-动作-提示四元组数据;其次,提出长时程一致性训练框架:通过建模预测残差并在训练中重注入不完美生成帧,使基础模型学会自我修正;同时,引入相机感知记忆检索与注入机制,实现长时程时空一致性;第三,设计基于分布匹配蒸馏(DMD)的多段自回归蒸馏策略,结合模型量化与VAE解码器剪枝,实现高效实时推理。实验表明,Matrix-Game 3.0在5B模型下可实现720p分辨率下最高40 FPS的实时生成,且保持分钟级序列的稳定记忆一致性;扩展至2×14B模型后,生成质量、动态表现与泛化能力进一步提升。该方法为工业级可部署世界模型提供了可行路径。
原文摘要 · Abstract (English)
With the advancement of interactive video generation, diffusion models have increasingly demonstrated their potential as world models. However, existing approaches still struggle to simultaneously achieve memory-enabled long-term temporal consistency and high-resolution real-time generation, limiting their applicability in real-world scenarios. To address this, we present Matrix-Game 3.0, a memory-augmented interactive world model designed for 720p real-time longform video generation. Building upon Matrix-Game 2.0, we introduce systematic improvements across data, model, and inference. First, we develop an upgraded industrial-scale infinite data engine that integrates Unreal Engine-based synthetic data, large-scale automated collection from AAA games, and real-world video augmentation to produce high-quality Video-Pose-Action-Prompt quadruplet data at scale. Second, we propose a training framework for long-horizon consistency: by modeling prediction residuals and re-injecting imperfect generated frames during training, the base model learns self-correction; meanwhile, camera-aware memory retrieval and injection enable the base model to achieve long horizon spatiotemporal consistency. Third, we design a multi-segment autoregressive distillation strategy based on Distribution Matching Distillation (DMD), combined with model quantization and VAE decoder pruning, to achieve efficient real-time inference. Experimental results show that Matrix-Game 3.0 achieves up to 40 FPS real-time generation at 720p resolution with a 5B model, while maintaining stable memory consistency over minute-long sequences. Scaling up to a 2x14B model further improves generation quality, dynamics, and generalization. Our approach provides a practical pathway toward industrial-scale deployable world models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。