arXiv:2410.07877cs.RO2024-10被引 7

无监督强化学习让四足机器人学会稳定自适应行走

Constrained Skill Discovery: Quadruped Locomotion with Unsupervised Reinforcement Learning

  • 通过最大化技能与状态的互信息并施加距离约束,学习可迁移的潜在表征
  • 相比基线方法,状态空间覆盖更广,运动行为更稳定可控
  • 无需任务奖励,在真实机器人上实现零样本任意位置移动

表示学习与无监督技能发现可使机器人在无需特定任务奖励的情况下习得多样且可复用的行为。本文采用无监督强化学习,通过最大化技能与状态之间的互信息,并施加距离约束来学习潜在表征。与先前受约束的技能发现方法相比,本方法以范数匹配目标替代潜转移最大化,不仅显著提升状态空间覆盖率,还使机器人学会更稳定、更易控制的行走行为。我们在真实ANYmal四足机器人上成功部署了所学策略,仅使用内在技能发现和标准正则化奖励,即可实现零样本下对笛卡尔状态空间中任意点的精确到达。

原文摘要 · Abstract (English)

Representation learning and unsupervised skill discovery can allow robots to acquire diverse and reusable behaviors without the need for task-specific rewards. In this work, we use unsupervised reinforcement learning to learn a latent representation by maximizing the mutual information between skills and states subject to a distance constraint. Our method improves upon prior constrained skill discovery methods by replacing the latent transition maximization with a norm-matching objective. This not only results in a much a richer state space coverage compared to baseline methods, but allows the robot to learn more stable and easily controllable locomotive behaviors. We successfully deploy the learned policy on a real ANYmal quadruped robot and demonstrate that the robot can accurately reach arbitrary points of the Cartesian state space in a zero-shot manner, using only an intrinsic skill discovery and standard regularization rewards.

四足机器人无监督学习技能发现强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。