arXiv:2603.08619cs.RO2026-03被引 4

将经典平衡控制原理融入强化学习,提升人形机器人抗跌能力

Embedding Classical Balance Control Principles in Reinforcement Learning for Humanoid Recovery

  • 用质心、捕捉点等物理指标做奖励和监督信号,指导策略学习
  • 在模拟中实现93.4%的恢复率,覆盖从脚踝到多接触站立的全场景
  • 无需预设动作轨迹,可直接部署到真实硬件,适合自主救援类任务

人形机器人在非结构化环境中仍易跌倒且难以恢复,限制了其实际应用。尽管强化学习已能实现起身动作,但现有方法将恢复视为纯任务奖励问题,未显式建模平衡状态。本文提出统一的强化学习策略,通过将经典平衡指标(捕捉点、质心状态、质心动量)作为特权评论器输入和奖励设计基础,在训练中直接围绕这些量构建奖励信号,而策略网络仅依赖本体感知以实现零样本硬件迁移。该策略无需参考轨迹或脚本接触,可覆盖完整恢复路径:小扰动时使用踝部与髋部策略,大推力下采用纠正步态,大范围跌倒时则通过手、肘、膝多接触实现柔韧坠落与起身。在Isaac Lab中的Unitree H1-2上训练,该策略在随机初始姿态与未预设跌倒配置下达到93.4%恢复率。消融实验表明,移除平衡引导结构会导致起身学习完全失败,验证了这些指标提供的是有意义的学习信号而非偶然结构。仿真到仿真迁移至MuJoCo及初步硬件实验进一步证明跨环境泛化能力。结果表明,将可解释的平衡结构嵌入学习框架,显著减少故障状态持续时间,并扩展自主恢复的边界。

原文摘要 · Abstract (English)

Humanoid robots remain vulnerable to falls and unrecoverable failure states, limiting their practical utility in unstructured environments. While reinforcement learning has demonstrated stand-up behaviors, existing approaches treat recovery as a pure task-reward problem without an explicit representation of the balance state. We present a unified RL policy that addresses this limitation by embedding classical balance metrics: capture point, center-of-mass state, and centroidal momentum, as privileged critic inputs and shaping rewards directly around these quantities during training, while the actor relies solely on proprioception for zero-shot hardware transfer. Without reference trajectories or scripted contacts, a single policy spans the full recovery spectrum: ankle and hip strategies for small disturbances, corrective stepping under large pushes, and compliant falling with multi-contact stand-up using the hands, elbows, and knees. Trained on the Unitree H1-2 in Isaac Lab, the policy achieves a 93.4% recovery rate across randomized initial poses and unscripted fall configurations. An ablation study shows that removing the balance-informed structure causes stand-up learning to fail entirely, confirming that these metrics provide a meaningful learning signal rather than incidental structure. Sim-to-sim transfer to MuJoCo and preliminary hardware experiments further demonstrate cross-environment generalization. These results show that embedding interpretable balance structure into the learning framework substantially reduces time spent in failure states and broadens the envelope of autonomous recovery.

人形机器人强化学习平衡控制自主恢复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。