arXiv:2409.15780cs.RO2024-09ICRA被引 22

用约束奖励让机器人学会多种走路方式,无需视觉传感器。

A Learning Framework for Diverse Legged Robot Locomotion Using Barrier-Based Style Rewards

  • 用松弛对数屏障函数设计风格奖励,引导学习期望步态
  • 45公斤机器人实现三足、两足行走,最高速度达4.67米/秒
  • 适合做复杂地形移动的足式机器人研究者参考

本文提出一种无模型强化学习框架,使足式机器人能实现多种运动模式(四足、三足或两足)及多样化任务。通过基于松弛对数屏障函数的运动风格奖励作为软约束,引导学习过程偏向目标运动风格,如步态、脚部抬升高度、关节位置或躯干高度。预设步态周期以灵活方式编码,支持学习过程中动态调整。大量实验表明,KAIST HOUND(45公斤机器人系统)可使用该框架实现两足、三足和四足行走;四足能力包括在不平坦地形行进、以4.67米/秒速度奔跑,以及跨越高达58厘米的障碍物(HOUND2可达67厘米);两足能力包括以3.6米/秒速度跑步、携带7.5公斤重物,以及攀爬楼梯——所有动作均无需外部感知输入。

原文摘要 · Abstract (English)

This work introduces a model-free reinforcement learning framework that enables various modes of motion (quadruped, tripod, or biped) and diverse tasks for legged robot locomotion. We employ a motion-style reward based on a relaxed logarithmic barrier function as a soft constraint, to bias the learning process toward the desired motion style, such as gait, foot clearance, joint position, or body height. The predefined gait cycle is encoded in a flexible manner, facilitating gait adjustments throughout the learning process. Extensive experiments demonstrate that KAIST HOUND, a 45 kg robotic system, can achieve biped, tripod, and quadruped locomotion using the proposed framework; quadrupedal capabilities include traversing uneven terrain, galloping at 4.67 m/s, and overcoming obstacles up to 58 cm (67 cm for HOUND2); bipedal capabilities include running at 3.6 m/s, carrying a 7.5 kg object, and ascending stairs-all performed without exteroceptive input.

足式机器人强化学习步态生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。