8分钟实机训练,实现四足机器人全向运动的高效学习
Gait in Eight: Efficient On-Robot Learning for Omnidirectional Quadruped Locomotion
- 采用CrossQ算法,极低计算开销下实现高效样本利用
- 8分钟真实时间训练完成全向步态学习,支持高速与稳定步态
- 适用于复杂室内外环境,适合需要快速部署的机器人研发
在机器人上进行强化学习是训练具备本体感知能力的腿式机器人策略的有前景方法。然而,实时学习对机器人计算资源的限制带来显著挑战。本文提出一种高效框架,仅需8分钟原始实时训练时间,便能利用新型非策略算法CrossQ的高样本效率和极低计算开销,实现四足机器人的全向运动学习。我们研究了两种控制架构:预测关节目标位置以实现敏捷高速运动,以及中央模式发生器用于生成稳定自然步态。与以往聚焦于简单前向步态的研究不同,本框架将机器人上的学习扩展至全向运动。我们在多种室内外环境中验证了该方法的鲁棒性。
原文摘要 · Abstract (English)
On-robot Reinforcement Learning is a promising approach to train embodiment-aware policies for legged robots. However, the computational constraints of real-time learning on robots pose a significant challenge. We present a framework for efficiently learning quadruped locomotion in just 8 minutes of raw real-time training utilizing the sample efficiency and minimal computational overhead of the new off-policy algorithm CrossQ. We investigate two control architectures: Predicting joint target positions for agile, high-speed locomotion and Central Pattern Generators for stable, natural gaits. While prior work focused on learning simple forward gaits, our framework extends on-robot learning to omnidirectional locomotion. We demonstrate the robustness of our approach in different indoor and outdoor environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。