arXiv:2409.13867cs.ROcs.AI2024-09被引 9

提出新算法让机器人安全控制自动收敛,提升可靠性与性能。

MAGICS: Adversarial RL with Minimax Actors Guided by Implicit Critic Stackelberg for Convergent Neural Synthesis of Robot Safety

  • 用极小极大对抗机制设计安全控制策略
  • 在模拟和真实四足机器人上均优于现有方法
  • 适合需要可证明安全性的机器人控制系统

尽管鲁棒最优控制理论能提供严格的安全控制策略,但在高维问题上难以扩展,导致深度学习被广泛用于可计算的安全合成。然而,现有神经安全合成方法通常缺乏收敛性保证与解的可解释性。本文提出最小极大对抗者协同隐式评判者斯塔克尔伯格框架(MAGICS),一种新型对抗强化学习算法,可保证局部收敛至极小极大均衡解。在此基础上,我们进一步为基于深度强化学习的通用机器人安全合成算法提供了局部收敛性保障。通过在OpenAI Gym环境中的仿真研究以及36维四足机器人的硬件实验,结果表明MAGICS能够生成优于当前最先进神经安全合成方法的鲁棒控制策略。

原文摘要 · Abstract (English)

While robust optimal control theory provides a rigorous framework to compute robot control policies that are provably safe, it struggles to scale to high-dimensional problems, leading to increased use of deep learning for tractable synthesis of robot safety. Unfortunately, existing neural safety synthesis methods often lack convergence guarantees and solution interpretability. In this paper, we present Minimax Actors Guided by Implicit Critic Stackelberg (MAGICS), a novel adversarial reinforcement learning (RL) algorithm that guarantees local convergence to a minimax equilibrium solution. We then build on this approach to provide local convergence guarantees for a general deep RL-based robot safety synthesis algorithm. Through both simulation studies on OpenAI Gym environments and hardware experiments with a 36-dimensional quadruped robot, we show that MAGICS can yield robust control policies outperforming the state-of-the-art neural safety synthesis methods.

强化学习机器人安全对抗训练收敛性保障

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。