让四足机器人在复杂环境中又快又安全地导航,训练时间仅需几分钟。
SEA-Nav: Efficient Policy Learning for Safe and Agile Quadruped Navigation in Cluttered Environments
- 用可微分屏障函数约束策略输出安全速度指令。
- 通过碰撞回放和危险探索奖励,提升关键经验学习效率。
- 支持真实世界部署,训练时间低至分钟级,适合工业落地。
在密集障碍物环境中高效训练四足机器人导航仍具挑战。现有方法或在简单障碍分布中缺乏安全性与敏捷性,或在复杂环境中运动缓慢,常需极长训练周期。为此,我们提出SEA-Nav(安全、高效、敏捷导航)强化学习框架。在多样且密集的障碍环境中,基于可微分控制屏障函数(CBF)的防护机制约束导航策略,确保输出安全速度指令。引入自适应碰撞回放机制与危险探索奖励,提升从关键经验中学习的概率,促进高效探索与利用。最后,融入运动学动作约束以保障速度指令安全,助力成功物理部署。据我们所知,这是首个在真实世界实现高度挑战性四足机器人导航且训练时间仅需分钟级的方法。
原文摘要 · Abstract (English)
Efficiently training quadruped robot navigation in densely cluttered environments remains a significant challenge. Existing methods are either limited by a lack of safety and agility in simple obstacle distributions or suffer from slow locomotion in complex environments, often requiring excessively long training phases. To this end, we propose SEA-Nav (Safe, Efficient, and Agile Navigation), a reinforcement learning framework for quadruped navigation. Within diverse and dense obstacle environments, a differentiable control barrier function (CBF)-based shield constraints the navigation policy to output safe velocity commands. An adaptive collision replay mechanism and hazardous exploration rewards are introduced to increase the probability of learning from critical experiences, guiding efficient exploration and exploitation. Finally, kinematic action constraints are incorporated to ensure safe velocity commands, facilitating successful physical deployment. To the best of our knowledge, this is the first approach that achieves highly challenging quadruped navigation in the real world with minute-level training time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。