构建可模拟多人交互的Minecraft世界模型,突破单视角局限。
Solaris: Building a Multiplayer Video World Model in Minecraft

- 设计支持多智能体协同的自动化数据采集系统
- 收集1264万帧多人游戏数据,验证多视角一致性
- 提出分阶段训练与检查点自强化,提升长程建模能力
现有动作条件视频生成模型(视频世界模型)局限于单智能体视角,无法捕捉真实环境中的多智能体交互。我们提出Solaris,一个可模拟一致多视角观测的多人视频世界模型。为实现这一目标,我们开发了专为Minecraft等游戏设计的多人数据系统,支持鲁棒、连续且自动化的数据采集。该系统不同于以往单人设置平台,能实现多智能体协同与同步视频+动作记录。基于此系统,我们收集了1264万帧多人游戏数据,并提出评估框架,涵盖多人移动、记忆、定位、建造和视图一致性。通过分阶段训练流程,逐步从单人建模过渡到多人建模,结合双向、因果与自强制训练。最终阶段引入检查点自强制(Checkpointed Self Forcing),一种内存高效的自强制变体,支持更长时序教师信号。结果表明,该架构与训练设计优于现有基线。通过开源系统与模型,我们希望为新一代多智能体世界模型奠定基础。
原文摘要 · Abstract (English)
Existing action-conditioned video generation models (video world models) are limited to single-agent perspectives, failing to capture the multi-agent interactions of real-world environments. We introduce Solaris, a multiplayer video world model that simulates consistent multi-view observations. To enable this, we develop a multiplayer data system designed for robust, continuous, and automated data collection on video games such as Minecraft. Unlike prior platforms built for single-player settings, our system supports coordinated multi-agent interaction and synchronized videos + actions capture. Using this system, we collect 12.64 million multiplayer frames and propose an evaluation framework for multiplayer movement, memory, grounding, building, and view consistency. We train Solaris using a staged pipeline that progressively transitions from single-player to multiplayer modeling, combining bidirectional, causal, and Self Forcing training. In the final stage, we introduce Checkpointed Self Forcing, a memory-efficient Self Forcing variant that enables a longer-horizon teacher. Results show our architecture and training design outperform existing baselines. Through open-sourcing our system and models, we hope to lay the groundwork for a new generation of multi-agent world models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。