开源模块化框架让世界-动作模型可实验、可复现,性能跨虚拟与现实保持领先。
OpenWAM: An Open, Modular Exploration Towards Systematic World-Action Model Pretraining
- 将世界-动作模型拆解为可组合模块,实现可控实验设计
- 在8个仿真与真实机器人任务中表现顶尖,跨领域泛化能力强
- 适合研究多模态预训练、具身智能与机器人控制的学者和开发者
世界-动作模型(WAM)从视频生成先验中继承世界知识,并通过具身经验转化为可执行控制信号。现有系统高度耦合,难以区分关键设计因素。我们提出OpenWAM,一个开放的研究栈,将世界-动作预训练转化为可控实验。OpenWAM-Infra将设计空间分解为统一训练、推理、部署与评估的可组合模块。基于此,OpenWAM-Study通过受控实验回答三个问题:继承什么、世界与动作学习如何交互、其协同效应如何扩展;并提炼出三条原则:上游知识通过足够强的生成骨干和紧凑信息丰富的隐空间传递;世界-动作协同需专用动作容量、显式世界到动作的信息流及同步联合去噪;具身预训练主要提升跨域泛化能力,单阶段联合训练人类与机器人数据可融合世界覆盖与动作锚定。据此构建OpenWAM-α,使用约6,400小时第一人称人类与机器人数据预训练,在八个仿真基准与真实机器人实验中表现优异,从仿真到物理世界均保持顶级性能。我们发布完整栈,包括基础设施、评估协议、预训练模型与数据配方,推动后续研究。
原文摘要 · Abstract (English)
World-Action Models inherit world knowledge from video-generative priors, and channel it into executable control signals through embodied experience. Existing systems, however, are monolithic: the generative backbone, visual representation, architecture, information flow, inference procedure, and training data are tightly coupled, obscuring which design choices matter and why. We introduce OpenWAM, an open research stack that turns world-action pretraining into a controlled experimental program. OpenWAM-Infra factorizes the WAM design space into composable modules with unified training, inference, deployment, and evaluation. On this substrate, OpenWAM-Study examines three questions through controlled experiments: what to inherit, how world and action learning interact, and how their synergy scales; and distills three principles: upstream knowledge transfers through a sufficiently capable generative backbone and a compact, information-rich latent space; world-action synergy requires dedicated action capacity, explicit world-to-action information flow, and synchronized joint denoising; and embodied pretraining principally improves out-of-domain generalization, with one-stage co-training over egocentric and robot data integrating world coverage and action grounding. Composing these principles, we build OpenWAM-{\alpha}, an open WAM pretrained on roughly 6,400 hours of egocentric human and robot data and evaluated across simulation and real-world benchmarks. Across the eight simulation benchmarks and the real-robot experiments, which together span embodiments from single-arm and bimanual manipulation to dexterous hands, OpenWAM-{\alpha} delivers consistently excellent performance, sustaining its top-tier standing from simulation to the physical world. We release the full stack, including infrastructure, evaluation protocols, pretrained models, and data recipes, to facilitate future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。