一个可提示的通用机器人控制模型,无需重训即可实现多种动作任务。
BFM-Zero: A Promptable Behavioral Foundation Model for Humanoid Control Using Unsupervised Reinforcement Learning
- 通过无监督强化学习构建统一动作-目标-奖励潜空间,支持多任务提示控制。
- 在真实Unitree G1机器人上实现零样本运动追踪、目标到达和奖励优化等全身体技能。
- 适合希望快速部署通用机器人控制策略的研究者与工程师。
为类人机器人构建行为基础模型(BFM)有望以单一可提示的通用策略统一多种控制任务。然而,现有方法要么仅限于仿真环境,要么专用于特定任务(如追踪)。本文提出BFM-Zero,一种将动作、目标与奖励嵌入统一潜空间的框架,使单个策略可通过提示完成多种下游任务而无需重训。该结构良好的潜空间使真实世界中的Unitree G1类人机器人实现了多样化的全身技能,包括零样本运动追踪、目标到达与奖励优化,并支持少样本优化适应。不同于以往基于策略的强化学习框架,BFM-Zero依托无监督强化学习与前向-后向(FB)模型,提供目标导向、可解释且平滑的全身动作潜表示。我们进一步引入关键奖励设计、域随机化及历史依赖的非对称学习以缩小仿真到现实的差距。关键设计在仿真中进行了量化消融实验。作为首个此类模型,BFM-Zero为可扩展的、可提示的行为基础模型在全身体控中的应用迈出重要一步。
原文摘要 · Abstract (English)
Building Behavioral Foundation Models (BFMs) for humanoid robots has the potential to unify diverse control tasks under a single, promptable generalist policy. However, existing approaches are either exclusively deployed on simulated humanoid characters, or specialized to specific tasks such as tracking. We propose BFM-Zero, a framework that learns an effective shared latent representation that embeds motions, goals, and rewards into a common space, enabling a single policy to be prompted for multiple downstream tasks without retraining. This well-structured latent space in BFM-Zero enables versatile and robust whole-body skills on a Unitree G1 humanoid in the real world, via diverse inference methods, including zero-shot motion tracking, goal reaching, and reward optimization, and few-shot optimization-based adaptation. Unlike prior on-policy reinforcement learning (RL) frameworks, BFM-Zero builds upon recent advancements in unsupervised RL and Forward-Backward (FB) models, which offer an objective-centric, explainable, and smooth latent representation of whole-body motions. We further extend BFM-Zero with critical reward shaping, domain randomization, and history-dependent asymmetric learning to bridge the sim-to-real gap. Those key design choices are quantitatively ablated in simulation. A first-of-its-kind model, BFM-Zero establishes a step toward scalable, promptable behavioral foundation models for whole-body humanoid control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。