arXiv:2411.13983cs.MAcs.RO2024-11中稿 · Proceeding at 2025…被引 4

用博弈论学习交互策略,让两车自主避让并协作通行

Learning Two-agent Motion Planning Strategies from Generalized Nash Equilibrium for Model Predictive Control

  • 通过求解广义纳什均衡生成数据,训练神经网络预测交互结果
  • 将预测结果作为终端代价函数,在MPC中实现自适应路径规划
  • 适用于竞速、无信号交叉口等复杂交互场景,适合自动驾驶研究

我们提出一种隐式博弈论模型预测控制(IGT-MPC),用于两智能体运动规划。该方法将博弈论中的交互结果预测值作为模型预测控制的终端代价函数,使智能体在不直接通信的情况下隐式考虑彼此影响并最大化自身收益。针对约束动态博弈问题,随机采样初始状态并求解广义纳什均衡(GNE),生成包含每轮交互收益的数据集;利用该数据训练简单神经网络以预测收益,进而构建终端代价函数。实验表明,IGT-MPC在两车对向竞速和无信号交叉口导航等场景中展现出竞争与协同行为。该方法融合机器学习与博弈推理,为基于模型的分布式多智能体运动规划提供新思路。

原文摘要 · Abstract (English)

We introduce an Implicit Game-Theoretic MPC (IGT-MPC), a decentralized algorithm for two-agent motion planning that uses a learned value function that predicts the game-theoretic interaction outcomes as the terminal cost-to-go function in a model predictive control (MPC) framework, guiding agents to implicitly account for interactions with other agents and maximize their reward. This approach applies to competitive and cooperative multi-agent motion planning problems which we formulate as constrained dynamic games. Given a constrained dynamic game, we randomly sample initial conditions and solve for the generalized Nash equilibrium (GNE) to generate a dataset of GNE solutions, computing the reward outcome of each game-theoretic interaction from the GNE. The data is used to train a simple neural network to predict the reward outcome, which we use as the terminal cost-to-go function in an MPC scheme. We showcase emerging competitive and coordinated behaviors using IGT-MPC in scenarios such as two-vehicle head-to-head racing and un-signalized intersection navigation. IGT-MPC offers a novel method integrating machine learning and game-theoretic reasoning into model-based decentralized multi-agent motion planning.

多智能体博弈论运动规划MPC

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。