用行动中心设计提升机器人控制效率,推理仅需85毫秒。
GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch

- 训练时用视觉动态辅助,推理时只生成动作,大幅减负。
- 在RTX 4090上实现85毫秒推理延迟,适合实时闭环部署。
- 通过自动搜索优化训练配置,减少调参成本。
世界动作模型(WAMs)通过联合建模动作与未来视觉观测,利用场景演化作为密集监督信号,提升机器人策略学习效果。然而,现有WAM普遍在推理时显式生成未来视频,带来巨大计算开销,阻碍实时闭环应用。GigaWorld-Policy提出以动作为中心的范式:训练阶段使用未来视觉动态,推理阶段仅生成动作。基于此框架,本文推出GigaWorld-Policy-0.5,采用混合动作条件世界建模(AC-WM)与WAM训练策略,增强视觉动态与动作的耦合性,提升动作表征迁移能力。为实现高效推理,引入分治式Mixture-of-Transformers架构,将视觉建模与动作生成拆分为专用专家,降低主动计算量,实现在本地RTX 4090上的85毫秒推理延迟。同时,采用基于智能体的AutoResearch流水线,系统化搜索训练配置,显著减少超参数调优的时间与人工干预。实验与消融验证表明,该模型在保留未来视觉动态训练优势的同时,大幅提升推理效率,适用于高效机器人控制。
原文摘要 · Abstract (English)
World Action Models (WAMs) improve robot policy learning by jointly modeling actions and future visual observations, using future scene evolution as dense supervision for physically grounded action generation. However, a common design in existing WAMs is to explicitly generate future videos at inference time, incurring substantial computational overhead and hindering real-time closed-loop deployment. GigaWorld-Policy addresses this issue with an action-centered formulation, where future visual dynamics are used during training while action-only decoding is used at inference time. Building upon this framework, we present GigaWorld-Policy-0.5, an enhanced action-centered WAM designed for more efficient robot control. During pretraining, GigaWorld-Policy-0.5 adopts a mixed Action-Conditioned World Modeling (AC-WM) and WAM training strategy. This strengthens the coupling between visual dynamics and robot actions and improves the transferability of action representations for downstream policy learning. For efficient inference, GigaWorld-Policy-0.5 introduces a Mixture-of-Transformers architecture that separates visual dynamics modeling and action generation into specialized experts, reducing active computation during action-only inference and achieving 85 ms inference latency on a local RTX 4090 setup. In addition, we employ an agent-based AutoResearch pipeline to systematically search training configurations, enabling more efficient identification of optimal experimental setups while reducing the time and manual intervention required for hyperparameter tuning. Experiments and ablations show that GigaWorld-Policy-0.5 preserves the training benefits of future visual dynamics while improving inference efficiency for robot control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。