arXiv:2510.18845cs.ROcs.SY2025-10被引 3

用MPC引导的深度学习,让机器人在对抗环境中更安全地决策。

MADR: MPC-guided Adversarial DeepReach

  • 结合模型预测控制与对抗性深度学习,求解双人零和微分博弈。
  • 在高维仿真与真实机器人上,优于现有最先进方法。
  • 适合需要鲁棒安全策略的复杂动态系统,如自动驾驶、机器人控制。

哈密顿-雅可比可达性提供了在对抗扰动下生成安全价值函数和策略的框架,但受限于维度诅咒。物理信息深度学习能够克服这一不可行性,但自身存在收敛缓慢且不准确的问题,主要源于弱偏微分方程梯度和自监督学习的复杂性。近期一些工作表明,通过引入基于最优控制问题本质的正则监督,能显著加速收敛并提升解的质量,然而这些方法仅限于单玩家问题和简单博弈。本文提出MADR:MPC引导的对抗性DeepReach,一个通用框架,用于稳健逼近双人零和微分博弈的价值函数。MADR不仅得到双方的最优策略,还能生成最坏情况下的安全策略。我们在多种高维仿真及真实机器人代理上测试该方法,涵盖不同动力学与博弈场景,结果表明其在仿真中显著优于现有最先进基线,在硬件上也取得了令人印象深刻的效果。

原文摘要 · Abstract (English)

Hamilton-Jacobi (HJ) Reachability offers a framework for generating safe value functions and policies in the face of adversarial disturbance, but is limited by the curse of dimensionality. Physics-informed deep learning is able to overcome this infeasibility, but itself suffers from slow and inaccurate convergence, primarily due to weak PDE gradients and the complexity of self-supervised learning. A few works, recently, have demonstrated that enriching the self-supervision process with regular supervision (based on the nature of the optimal control problem), greatly accelerates convergence and solution quality, however, these have been limited to single player problems and simple games. In this work, we introduce MADR: MPC-guided Adversarial DeepReach, a general framework to robustly approximate the two-player, zero-sum differential game value function. In doing so, MADR yields the corresponding optimal strategies for both players in zero-sum games as well as safe policies for worst-case robustness. We test MADR on a multitude of high-dimensional simulated and real robotic agents with varying dynamics and games, finding that our approach significantly out-performs state-of-the-art baselines in simulation and produces impressive results in hardware.

博弈论机器人安全深度学习控制优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。