arXiv:2504.17128eess.SYcs.LG2025-04中稿 · 7th Annual Confere…被引 9

提出实时推断对手目标函数的控制框架,解决信息不全下的博弈问题。

PACE: A Framework for Learning and Control in Linear Incomplete-Information Differential Games

  • 将对手视为学习者,建模其动态以推断成本参数
  • 理论证明参数估计收敛且系统状态稳定
  • 适用于人机交互、多智能体控制等场景

本文研究了两玩家线性二次微分博弈中的不完全信息问题,常见于多智能体控制、人-机器人交互(HRI)及一般和博弈的近似求解。传统方法依赖耦合Riccati方程,但当双方均不知对方代价函数时,求解复杂度显著上升。为此,我们提出基于模型的同伴感知成本估计(PACE)框架,使每个智能体将对方视为学习者,建模其学习动态,并利用该动态实时推断对方的代价函数参数。该方法仅需前序状态观测即可实时推理对方目标,并动态调整自身控制策略。此外,我们提供了参数估计收敛性和系统状态稳定性的理论保证。数值实验表明,相较于将对方近似为具有完全信息的静态最优策略,建模对方学习动态能显著提升稳定性与收敛速度。

原文摘要 · Abstract (English)

In this paper, we address the problem of a two-player linear quadratic differential game with incomplete information, a scenario commonly encountered in multi-agent control, human-robot interaction (HRI), and approximation methods for solving general-sum differential games. While solutions to such linear differential games are typically obtained through coupled Riccati equations, the complexity increases when agents have incomplete information, particularly when neither is aware of the other's cost function. To tackle this challenge, we propose a model-based Peer-Aware Cost Estimation (PACE) framework for learning the cost parameters of the other agent. In PACE, each agent treats its peer as a learning agent rather than a stationary optimal agent, models their learning dynamics, and leverages this dynamic to infer the cost function parameters of the other agent. This approach enables agents to infer each other's objective function in real time based solely on their previous state observations and dynamically adapt their control policies. Furthermore, we provide a theoretical guarantee for the convergence of parameter estimation and the stability of system states in PACE. Additionally, in our numerical studies, we demonstrate how modeling the learning dynamics of the other agent benefits PACE, compared to approaches that approximate the other agent as having complete information, particularly in terms of stability and convergence speed.

博弈控制在线学习多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。