让智能体实时推断对手目标并协同学习,提升多智能体协作效率
Peer-Aware Cost Estimation in Nonlinear General-Sum Dynamic Games for Mutual Learning and Intent Inference
- 基于迭代线性二次近似,双智能体互相建模对方学习过程
- 实测显示对未知目标函数的推断速度和精度显著优于基线方法
- 适用于人机协作、自动驾驶等需动态理解意图的场景
动态博弈论是建模多智能体交互与人机系统的重要工具。现实中,双方目标函数常不完全透明,需建模为不完全信息的一般和动态博弈。此类博弈的均衡策略求解极具挑战,尤其当系统存在非线性动态时。现有方法常假设一方为全知专家,导致估计偏差与协作失败。为此,本文提出非线性同行感知代价估计(N-PACE)算法。N-PACE利用迭代线性二次(ILQ)近似,使每个智能体在实时更新自身控制策略的同时,显式建模对手的学习动态,推断其目标函数,实现无偏且快速的目标学习。此外,通过显式建模对手学习行为,N-PACE支持意图通信。理论分析与案例研究均表明,相较于忽略对手学习行为的基线方法,N-PACE在性能上具有显著优势。
原文摘要 · Abstract (English)
Dynamic game theory is a powerful tool in modeling multi-agent interactions and human-robot systems. In practice, since the objective functions of both agents may not be explicitly known to each other, these interactions can be modeled as incomplete-information general-sum dynamic games. Solving for equilibrium policies for such games presents a major challenge, especially if the games involve nonlinear underlying dynamics. To simplify the problem, existing work often assumes that one agent is an expert with complete information about its peer, which can lead to biased estimates and failures in coordination. To address this challenge, we propose a nonlinear peer-aware cost estimation (N-PACE) algorithm for general-sum dynamic games. In N-PACE, using iterative linear quadratic (ILQ) approximation of dynamic games, each agent explicitly models the learning dynamics of its peer agent while inferring their objective functions and updating its own control policy accordingly in real time, which leads to unbiased and fast learning of the unknown objective function of the peer agent. Additionally, we demonstrate how N-PACE enables intent communication by explicitly modeling the peer's learning dynamics. Finally, we show how N-PACE outperforms baseline methods that disregard the learning behavior of the other agent, both analytically and using our case studies
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。