统一多专家世界模型,让不同动作控制协同训练。
Worldscape-MoE: A Unified Mixture-of-Experts World Model for Scalable Heterogeneous Action Control

- 用多专家架构融合多种动作输入,共享世界动态表征。
- 在多场景任务中提升运动与手控表现,支持新动作持续扩展。
- 适合需要多模态动作控制的智能体研究与开发。
世界模型正成为具身智能与交互代理的核心基础设施,提供可控制的模拟环境以实现感知、行动、预测和可扩展的经验获取。然而,当前视频生成的世界模型仍围绕孤立的控制接口组织,如相机轨迹、机器人动作或手部关节信号,这种碎片化已成为扩展瓶颈。核心挑战并非缺乏可控生成器,而是缺乏一个能吸收异构动作监督并保持共享世界动态模型的统一可扩展学习框架。本文提出Worldscape-MoE,基于扩散Transformer的多专家世界模型,实现可扩展的异构动作控制。关键观察是:不同控制方式虽表示各异,但均约束同一世界的物理规律、场景动态与交互语义。通过模态感知的动作注入、共享与特定专家设计,以及渐进式多专家调优策略,该模型支持新动作模态的持续扩展。跨行走、机器人操作和第一人称手控实验表明,异构监督反而提升单个控制能力。Worldscape-MoE在WorldArena上表现优异,改进行走与手控指标,展现强分布外泛化能力,并随控制数据与专家增加呈现良好扩展性。
原文摘要 · Abstract (English)
World models are rapidly becoming a core infrastructure for embodied intelligence and interactive agents: they provide controllable simulators in which agents can perceive, act, forecast, and acquire scalable experience. Yet current video generation world models are still organized around isolated control interfaces, such as camera trajectories, robot actions, or hand-joint signals. This fragmentation is increasingly a scaling bottleneck. The central challenge is not the absence of controllable generators, but the lack of a unified and extensible learning framework that can absorb heterogeneous action supervision while preserving a shared model of world dynamics. In this work, we introduce Worldscape-MoE, a Mixture-of-Experts world model built on Diffusion Transformers for scalable heterogeneous action control. Our key observation is that different controls specify different interfaces to the same underlying world: although their representations differ, they constrain shared physical regularities, scene dynamics, and interaction semantics. Worldscape-MoE operationalizes this observation through modality-aware control injection, shared and control-specific experts, and a progressive MoE tuning strategy that supports continual extension to new action modalities. Experiments across locomotion, robotic manipulation, and egocentric hand control show that heterogeneous supervision improves rather than interferes with individual control capabilities. Worldscape-MoE achieves strong results on WorldArena, improves locomotion and hand-control metrics, exhibits robust out-of-distribution generalization, and demonstrates scaling behavior as additional control data and experts are integrated.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。