arXiv:2503.11467cs.ROcs.LG2025-03被引 2

用受限理性的对抗强化学习提升四足机器人避障能力。

Dynamic Obstacle Avoidance with Bounded Rationality Adversarial Reinforcement Learning

  • 将障碍物建模为理性受限的对抗智能体,通过分级策略训练导航
  • 在随机迷宫中实现95%以上成功率,优于传统方法
  • 适合真实机器人部署,尤其适用于复杂动态环境

强化学习在获取腿式机器人稳定步态方面已证明有效,但设计能在未知环境中稳健避障的控制算法仍是四足运动中的持续挑战。为此,采用分层方法——低层运动策略与高层导航策略结合。关键在于,高层策略需对路径上的动态障碍具有鲁棒性。本文提出一种新方法:通过对抗强化学习范式,将障碍物视为对抗智能体进行训练。为提升训练可靠性,引入量化响应均衡(quantal response equilibria)限制对抗智能体的理性程度,并施加理性程度渐进式课程。该方法称为基于量化响应的对抗强化学习(Hi-QARL)。我们在包含多个障碍物的随机未见迷宫中对其进行了基准测试,结果显示在复杂动态环境下成功率超过95%。为验证实际适用性,该方法在模拟的Unitree GO1机器人上成功应用。

原文摘要 · Abstract (English)

Reinforcement Learning (RL) has proven largely effective in obtaining stable locomotion gaits for legged robots. However, designing control algorithms which can robustly navigate unseen environments with obstacles remains an ongoing problem within quadruped locomotion. To tackle this, it is convenient to solve navigation tasks by means of a hierarchical approach with a low-level locomotion policy and a high-level navigation policy. Crucially, the high-level policy needs to be robust to dynamic obstacles along the path of the agent. In this work, we propose a novel way to endow navigation policies with robustness by a training process that models obstacles as adversarial agents, following the adversarial RL paradigm. Importantly, to improve the reliability of the training process, we bound the rationality of the adversarial agent resorting to quantal response equilibria, and place a curriculum over its rationality. We called this method Hierarchical policies via Quantal response Adversarial Reinforcement Learning (Hi-QARL). We demonstrate the robustness of our method by benchmarking it in unseen randomized mazes with multiple obstacles. To prove its applicability in real scenarios, our method is applied on a Unitree GO1 robot in simulation.

强化学习四足机器人避障对抗训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。