arXiv:2409.19746cs.RO2024-09ICRA被引 1

用可解释的最坏扰动训练鲁棒策略,提升机器人抗干扰能力。

Learning Robust Policies via Interpretable Hamilton-Jacobi Reachability-Guided Disturbances

  • 基于哈密顿-雅可比可达性生成可解释扰动作为对抗训练对手
  • 在仿真与真实环境中均实现稳定性能,批评网络与理论值函数一致
  • 适合关注机器人安全控制与对抗训练的研究者

深度强化学习在具有复杂异构动力学的机器人任务中表现优异,但对未知扰动和对抗攻击仍显脆弱。本文提出一种融合模型基础控制与对抗强化学习的鲁棒策略训练框架,无需外部黑盒对抗器。方法引入一种新型哈密顿-雅可比可达性引导的扰动,以可解释的最坏或近最坏情况扰动作为对抗训练中的对手。我们在三个任务中验证其有效性:仿真与真实世界中的避障任务,以及仿真中的四旋翼稳定控制任务。结果表明,所学批评网络与真实HJ值函数保持一致,策略网络性能与其他学习方法相当。

原文摘要 · Abstract (English)

Deep Reinforcement Learning (RL) has shown remarkable success in robotics with complex and heterogeneous dynamics. However, its vulnerability to unknown disturbances and adversarial attacks remains a significant challenge. In this paper, we propose a robust policy training framework that integrates model-based control principles with adversarial RL training to improve robustness without the need for external black-box adversaries. Our approach introduces a novel Hamilton-Jacobi reachability-guided disturbance for adversarial RL training, where we use interpretable worst-case or near-worst-case disturbances as adversaries against the robust policy. We evaluated its effectiveness across three distinct tasks: a reach-avoid game in both simulation and real-world settings, and a highly dynamic quadrotor stabilization task in simulation. We validate that our learned critic network is consistent with the ground-truth HJ value function, while the policy network shows comparable performance with other learning-based methods.

强化学习鲁棒控制对抗训练可达性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。