用引导点强化学习,让双臂机器人像猿猴一样稳健摆荡。
Robust Brachiation on a Life-Sized Dual-Arm Robot Using Waypoint-Guided Reinforcement Learning

- 通过稀疏引导点控制末端轨迹,结合强化学习生成复杂动作。
- 在仿真与真实机器人上均实现稳定摆荡及故障恢复能力。
- 适合研究仿人机器人运动控制与复杂环境移动的学者参考。
摆荡是一种灵长类动物主要依靠手臂进行移动的运动方式,可在无落脚点环境中通行。然而,该动作需要高度协调的整体运动和精确的抓握与释放时机控制,因此在全尺寸机器人平台上实现稳健行为仍具挑战。本研究提出一种基于强化学习的方法,实现在全尺寸双臂机器人上的摆荡。核心方法为路径点引导强化学习(WGRL),用于生成非线性复杂运动。对于缺乏模仿学习数据的高难度任务,该方法通过稀疏指定末端执行器轨迹的路径点,由强化学习生成全身运动。通过将路径点跟踪引导与任务成功奖励及机械能奖励相结合,并在支持从仿真到现实迁移的环境中训练,所提方法实现了前进推进与运动稳定性。在具有几何变化的猴爪杆环境中的仿真-仿真实验以及硬件实验中评估了所获行为,验证了其鲁棒性,包括故障恢复能力。本研究为在全尺寸机器人硬件上实现基于手臂的运动提供了有效的学习设计指南,并扩展了机器人的可通行工作空间。
原文摘要 · Abstract (English)
Brachiation is a form of locomotion in which primates move primarily using their arms, enabling traversal in environments without footholds. However, this motion requires highly coordinated whole-body movement and precise timing control for bar grasping and release. As a result, achieving robust behavior on life-sized robotic platforms remains challenging. In this study, we present a reinforcement learning-based method to realize brachiation on a life-sized dual-arm robot. The core of the proposed approach is Waypoint-Guided Reinforcement Learning (WGRL), a learning framework for inducing non-linear and complex motions. For high-difficulty tasks where imitation learning data are unavailable, WGRL guides behavior acquisition by sparsely specifying waypoints for the end-effector trajectory, while whole-body motion is generated through reinforcement learning. In addition, by integrating the waypoint-following guidance with rewards based on task success and mechanical energy, and training in an environment designed for Sim-to-Real transfer, the proposed method achieves both forward progression and motion stability. The acquired behavior is evaluated through Sim-to-Sim experiments under monkey-bar environments with geometric variations and hardware experiments, confirming robust brachiation including failure recovery behavior. This study provides effective learning design guidelines for realizing arm-based locomotion on life-sized robotic hardware and expanding the traversable workspace of robots.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。