arXiv:2503.07013cs.ROcs.GT2025-03中稿 · 2025 ACC

提出新方法,更高效学习双人避障博弈的均衡策略。

Learning Nash Equilibrial Hamiltonian for Two-Player Collision-Avoiding Interactions

  • 利用共状态结构简化学习,降低数据需求。
  • 在相同采样预算下,碰撞概率比现有方法更低。
  • 适合需要低样本开销的多智能体避障场景。

本文研究双人风险敏感型避障交互中的纳什均衡策略学习问题。由于均衡值在状态空间上存在不连续性,实时求解广义微分博弈的哈密顿-雅可比-伊斯克斯方程仍具挑战。现有方法通常依赖大量不同初始状态下的均衡策略样本进行监督训练以降低碰撞风险。本文提出两项改进:首先,在系统动力学线性且避障主导损失函数时,证明均衡共状态具有简单结构,更易高效学习;其次,引入基于理论的主动学习机制,通过庞特里亚金最大原理的符合度作为采集函数指导数据采样。在无控制交叉口案例中,该方法在相同数据采集预算下,实现了更泛化的均衡策略逼近,显著降低碰撞概率,优于当前最优方法。

原文摘要 · Abstract (English)

We consider the problem of learning Nash equilibrial policies for two-player risk-sensitive collision-avoiding interactions. Solving the Hamilton-Jacobi-Isaacs equations of such general-sum differential games in real time is an open challenge due to the discontinuity of equilibrium values on the state space. A common solution is to learn a neural network that approximates the equilibrium Hamiltonian for given system states and actions. The learning, however, is usually supervised and requires a large amount of sample equilibrium policies from different initial states in order to mitigate the risks of collisions. This paper claims two contributions towards more data-efficient learning of equilibrium policies: First, instead of computing Hamiltonian through a value network, we show that the equilibrium co-states have simple structures when collision avoidance dominates the agents' loss functions and system dynamics is linear, and therefore are more data-efficient to learn. Second, we introduce theory-driven active learning to guide data sampling, where the acquisition function measures the compliance of the predicted co-states to Pontryagin's Maximum Principle. On an uncontrolled intersection case, the proposed method leads to more generalizable approximation of the equilibrium policies, and in turn, lower collision probabilities, than the state-of-the-art under the same data acquisition budget.

博弈学习避障主动学习纳什均衡

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。