让机器人走路更自然,还能稳稳摔倒后站起来。
Predictive Style Matching: Natural and Robust Humanoid Locomotion

- 用历史状态预测上半身动作目标,训练时指导奖励
- 实测上半身动作误差降低近10倍,跌倒恢复率不变
- 适合需要自然姿态又兼顾鲁棒性的机器人控制
强化学习已成为人形机器人行走控制的主流方法:策略能可靠从仿真迁移到硬件,并在受扰后稳健恢复。然而运动质量仍不理想:仅基于任务的奖励常导致僵硬、不对称步态;而运动模仿方法虽改善外观,却因参考信号与平衡所需姿态冲突,对外部扰动更敏感。本文提出预测性风格匹配(Predictive Style Matching),通过离线预测器将机器人下肢状态历史和速度指令映射为可解释的上半身关节与步态目标,用于训练阶段的奖励设计。由于目标是状态相关而非时间依赖,且预测器仅用于训练,部署控制器保持了任务型强化学习基线的本体感知接口与推理开销。在Unitree G1机器人上,无论是仿真还是真实硬件测试中,PSM相比仅任务型强化学习,上半身风格误差降低约一个数量级,同时保持相同的跌倒恢复率;而运动模仿基线虽达到最低风格误差,但受扰后无法恢复的次数多出约五倍。
原文摘要 · Abstract (English)
Reinforcement learning has become the prevailing approach to humanoid locomotion control: policies transfer reliably from simulation to hardware and recover gracefully from disturbances. Motion quality, however, still lags behind: task-only rewards often converge to stiff, asymmetric gaits, while motion imitation methods improve appearance but become more sensitive to external disturbances because reference signals can oppose the transient poses needed to regain balance. We propose Predictive Style Matching, in which an offline predictor maps the robot's lower-body state history and velocity commands to interpretable upper-body joint and gait targets that shape the rewards during training. Because the targets are state-conditioned rather than time-indexed and the predictor is used only at training time, the deployed controller inherits the proprioceptive interface and inference cost of a task-only RL baseline. On the Unitree G1, in both simulation and hardware, PSM reduces upper-body style error by roughly an order of magnitude over task-only RL while preserving its fall-recovery rate, whereas the motion-imitation baseline attains the lowest style error but fails to recover from disturbances about five times as often.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。