让机器人在复杂环境中自主平衡,自动修正错误指令。
GigaBrain-WBC-0.5: A Behavior World Model for Robust Whole-Body Control with Environment Interaction

- 用因果Transformer联合预测动作、状态和行为分布,理解环境影响。
- 实测在地形交互中成功率81.3%(最强基线4.3倍),跌倒恢复率99.3%(16.8倍)。
- 适合需要鲁棒人形控制的科研与工程场景,支持真实部署。
全身运动追踪策略将人形机器人变为鲁棒控制接口:操作者或上游模型仅提供粗略运动意图,底层策略负责保持平衡与物理可行性。现有追踪器仅适用于平坦地面:在空场景中训练,未学习接触地形与物体如何重塑其动力学,且通过不断扩充参考动作数据集来应对各种命令,一旦可行行为依赖环境便失效。我们提出GigaBrain-WBC-0.5,首个用于人形机器人全身控制的行为世界模型(BWM)。不同于纯反应式追踪器,我们训练一个因果Transformer,联合预测下一步动作、状态及下一步潜在行为命令的分布,使执行网络同时建模环境对行为的塑造作用。自动化的地形标注流程从重定向运动中恢复完整3D接触几何,实现与现有动作数据集规模相当的标注。预测分布部署时用于在线检测不可行命令并将其回退至已学行为,使机器人以“尽力而为”方式执行任务。结果是一个统一策略:实时接收命令,与环境交互,并对不可行指令、跌倒和扰动保持鲁棒性。GigaBrain-WBC-0.5在四个场景中均优于三个大规模追踪基线:地形交互成功率81.3%(最强基线4.3倍),不可行指令下成功率83.1%,跌倒恢复成功率99.3%(最强基线16.8倍)。硬件测试表明其在支撑缺失与扰动下仍具鲁棒性;Unitree G1检查点经简单微调即可迁移到Maker L01机器人。
原文摘要 · Abstract (English)
Whole-body motion tracking policies turn a humanoid into a robust control interface: the teleoperator---or an upstream model---only supplies a coarse movement intent, while the low-level policy keeps the robot balanced and physically feasible. Existing trackers deliver this interface only on flat ground: trained in empty scenes, they never learn how contact with terrain and objects reshapes their dynamics, and they attempt to teach the policy to balance under any command by continually enlarging the reference-motion corpus, which stops working once feasible behaviors become environment-dependent. We present GigaBrain-WBC-0.5, the first Behavior World Model (BWM) for humanoid whole-body control. Rather than a purely reactive tracker, we train a causal Transformer to jointly predict its next action, next state, and the distribution over its next latent behavior command, so the network that acts also models how the environment shapes what it can do next. An automatic terrain-annotation pipeline recovers full 3D contact geometry from retargeted motion, enabling terrain annotation at the scale of existing motion datasets. The predicted distribution is reused at deployment to detect implausible commands online and retract them onto learned behaviors, so the robot attempts tasks in a "best-effort" manner. The result is a unified policy that takes real-time command, interacts with environment, and stays robust to implausible commands, falls, and disturbances. GigaBrain-WBC-0.5 achieves the highest success rate across all four regimes among three large-scale tracker baselines: 81.3% on terrain interaction (4.3x the strongest baseline), 83.1% under implausible commands, and 99.3% fall recovery (16.8x the strongest baseline). Hardware trials show robust interaction under missing supports and disturbances; the Unitree G1 checkpoint transfers to the Maker L01 robot with simple fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。