用无监督技能发现自动探索,让四足机器人学会攀爬跳跃等敏捷动作。
Unsupervised Skill Discovery as Exploration for Learning Agile Locomotion
- 通过无监督方法自动学习多样化技能,替代人工设计探索策略。
- 在模拟中实现爬行、跳跃、垂直墙跃起等复杂动作,无需奖励工程。
- 适合研究机器人自主学习与强化学习探索机制的研究者。
探索对腿式机器人学习克服多样障碍的敏捷运动行为至关重要。然而,这种探索本质上具有挑战性,通常依赖大量奖励工程、专家示范或课程学习,均限制了泛化能力。本文提出技能发现作为探索(SDAX),一种新颖的学习框架,显著减少人工干预。SDAX利用无监督技能发现,自主获取应对障碍的多样化技能组合;通过双层优化动态调节训练过程中的探索程度。实验表明,SDAX使四足机器人学会了包括爬行、攀爬、跳跃及从垂直墙面跃起在内的高度敏捷行为。最终,所学策略成功部署于真实硬件,验证了其向现实世界的有效迁移。
原文摘要 · Abstract (English)
Exploration is crucial for enabling legged robots to learn agile locomotion behaviors that can overcome diverse obstacles. However, such exploration is inherently challenging, and we often rely on extensive reward engineering, expert demonstrations, or curriculum learning - all of which limit generalizability. In this work, we propose Skill Discovery as Exploration (SDAX), a novel learning framework that significantly reduces human engineering effort. SDAX leverages unsupervised skill discovery to autonomously acquire a diverse repertoire of skills for overcoming obstacles. To dynamically regulate the level of exploration during training, SDAX employs a bi-level optimization process that autonomously adjusts the degree of exploration. We demonstrate that SDAX enables quadrupedal robots to acquire highly agile behaviors including crawling, climbing, leaping, and executing complex maneuvers such as jumping off vertical walls. Finally, we deploy the learned policy on real hardware, validating its successful transfer to the real world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。