让机器人走路更协调,模仿人体对称运动机制
Coordinated Humanoid Robot Locomotion with Symmetry Equivariant Reinforcement Learning Policy
- 在强化学习中引入对称性约束,确保左右动作一致
- 仿真与真实机器人测试中,追踪精度提升最高40%
- 适合需要稳定协调运动的类人机器人应用
人类神经系统具有双侧对称性,支持协调平衡的运动。但现有深度强化学习方法忽视机器人本体的形态对称性,导致动作不协调且性能不佳。受人体运动控制启发,本文提出对称等变策略(SE-Policy),在智能体中嵌入严格对称等变性,在评价器中引入对称不变性,无需额外超参数。该方法强制对称观测下行为一致,生成时空协调的动作,显著提升任务表现。在速度追踪任务中,通过仿真和真实世界部署(使用Unitree G1 humanoid机器人)的大量实验表明,相比最先进基线方法,SE-Policy将追踪精度提升最高达40%,并实现更优的空间-时间协调性。结果证明了该方法的有效性及其在类人机器人中的广泛适用性。
原文摘要 · Abstract (English)
The human nervous system exhibits bilateral symmetry, enabling coordinated and balanced movements. However, existing Deep Reinforcement Learning (DRL) methods for humanoid robots neglect morphological symmetry of the robot, leading to uncoordinated and suboptimal behaviors. Inspired by human motor control, we propose Symmetry Equivariant Policy (SE-Policy), a new DRL framework that embeds strict symmetry equivariance in the actor and symmetry invariance in the critic without additional hyperparameters. SE-Policy enforces consistent behaviors across symmetric observations, producing temporally and spatially coordinated motions with higher task performance. Extensive experiments on velocity tracking tasks, conducted in both simulation and real-world deployment with the Unitree G1 humanoid robot, demonstrate that SE-Policy improves tracking accuracy by up to 40% compared to state-of-the-art baselines, while achieving superior spatial-temporal coordination. These results demonstrate the effectiveness of SE-Policy and its broad applicability to humanoid robots.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。