不依赖模型蒸馏,实现人形机器人单腿平衡的可部署控制。
A Change of Frame Makes Balance Observable: Distillation-Free Humanoid Single-Leg Stance

- 将动态质心观测引入硬件控制器,通过足底坐标系消除不可测速度影响。
- 在90个测试动作中保持89个干净单腿平衡,真实机器人验证成功。
- 提供首个方法无关的仿真到仿真的基准,推动平衡能力可测量化。
统一的人形机器人策略能处理复杂全身运动,但在单腿站立这一基础任务上表现不佳。在我们的单腿平衡基准测试中,八种公开发布的先进通用策略在90个测试动作中均未能保持稳定,仅靠迈步或跳跃来恢复失衡,而非预防。预防失衡需用到捕获点(xCoM),即由质心速度外推得到的点,但因其依赖基座线速度,而该量无法由机载传感器直接测量,故从未被用于训练实际运行的策略。本文提出改变参考坐标系:以支撑脚为参考系,该速度恰好抵消,使观测仅通过编码器与惯性测量单元即可重建。我们将这一首个可部署的动态质心观测直接嵌入硬件运行的执行器中,并结合逐项翻译自人类姿势控制的奖励库,遵循“预防优于修复”原则。采用非对称FastSAC训练,无需蒸馏,所得策略DDC在九类分层姿态下,90个未见动作中保持89个干净单腿平衡,并成功迁移到真实Unitree G1机器人;消融实验表明,动态质心观测是性能提升的最大因素,移除它导致清洁单腿平衡得分下降43分。我们发布完整工具链,包含首个方法无关、可复现的模拟到模拟的单腿平衡基准,评估每个策略在与训练环境不同的模拟器中的表现,助力将平衡从特定任务技巧转变为可度量、可构建的核心能力。
原文摘要 · Abstract (English)
Unified humanoid policies handle agile whole-body motion, yet stumble on a simple demand: staying balanced on one leg. On our single-leg-balance benchmark, eight released state-of-the-art general policies hold a clean single-leg stance on 0 of 90 test motions; they stay up only by stepping or hopping, recovering from imbalance rather than preventing it. Prevention needs the capture point (xCoM), the center of mass (CoM) extrapolated by its velocity, which has never driven a learned hardware policy because it requires a base linear velocity that no on-board sensor measures directly. A change of frame makes it observable: expressed relative to the support foot, that velocity cancels exactly, leaving an observation reconstructible from encoders and IMU alone. We put this first deployable dynamic-CoM observation directly into the actor that runs on hardware, and pair it with a reward library translated term by term from human postural control, under one principle: prevention over repair. Trained via asymmetric FastSAC without distillation, the resulting policy, DDC (Deployable Dynamic-CoM), holds clean single-leg balance on 89 of 90 held-out motions across nine stratified pose classes and transfers to a real Unitree G1; in ablation, the dynamic-CoM observation is the single largest driver: removing it alone costs 43 points of clean single-leg balance. We release the full stack with the first method-agnostic, reproducible sim2sim benchmark for humanoid single-leg balance, scoring each policy in a simulator distinct from the one it was trained in, to help turn balance from a per-task trick into a capability the field can measure and build in.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。