让机器人像人一样摆臂,用分治强化学习提升行走平衡性
Learning Humanoid Arm Motion via Centroidal Momentum Regularized Multi-Agent Reinforcement Learning
- 分设手脚独立智能体,共享基底状态和质心角动量信息
- 摆臂动作减少整体角动量,使机器人在崎岖地形上更稳
- 适合做复杂运动控制的机器人研发者参考
人类在行走时自然摆臂以调节全身动力学、减小角动量并维持平衡。受此启发,我们提出一种基于肢体级多智能体强化学习的框架,通过涌现的臂部运动实现人形机器人的协调全身控制。该方法为四肢分别设计独立的演员-评论家结构,采用集中式评论家与去中心化演员,仅共享基底状态和质心角动量(CAM)观测,使各智能体可通过模块化奖励设计专注于特定任务。臂部智能体通过追踪和阻尼CAM的奖励机制,引导出降低整体角动量与垂直地面反作用力矩的摆臂行为,显著提升行走或受外力扰动时的稳定性。与单智能体及其它多智能体基线对比验证了本方法的有效性。最终,我们将学习到的策略部署于人形平台,在平坦路面行走、复杂地形穿越和楼梯攀爬等多样任务中均实现了鲁棒性能。
原文摘要 · Abstract (English)
Humans naturally swing their arms during locomotion to regulate whole-body dynamics, reduce angular momentum, and help maintain balance. Inspired by this principle, we present a limb-level multi-agent reinforcement learning (RL) framework that enables coordinated whole-body control of humanoid robots through emergent arm motion. Our approach employs separate actor-critic structures for the arms and legs, trained with centralized critics but decentralized actors that share only base states and centroidal angular momentum (CAM) observations, allowing each agent to specialize in task-relevant behaviors through modular reward design. The arm agent guided by CAM tracking and damping rewards promotes arm motions that reduce overall angular momentum and vertical ground reaction moments, contributing to improved balance during locomotion or under external perturbations. Comparative studies with single-agent and alternative multi-agent baselines further validate the effectiveness of our approach. Finally, we deploy the learned policy on a humanoid platform, achieving robust performance across diverse locomotion tasks, including flat-ground walking, rough terrain traversal, and stair climbing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。