通过反向传播识别动态成本,还原博弈中各方的隐含目标。
Identifying Time-varying Costs in Finite-horizon Linear Quadratic Gaussian Games
- 用反向传播从均衡策略中逆推时变成本参数
- 在有限轨迹下给出成本识别的概率误差边界
- 适用于仿真与驾驶场景,可复现真实轨迹
我们研究有限时域线性二次高斯博弈中的成本识别问题。刻画了生成给定纳什均衡策略的所有成本参数集合,提出一种反向传播算法以识别时变成本参数,并在仅基于有限轨迹的情况下推导出成本识别的的概率误差界。我们在数值模拟和驾驶模拟中验证了该方法的有效性。实验表明,该算法能准确识别出可重现纳什均衡策略及观测轨迹的成本参数。
原文摘要 · Abstract (English)
We address cost identification in a finite-horizon linear quadratic Gaussian game. We characterize the set of cost parameters that generate a given Nash equilibrium policy. We propose a backpropagation algorithm to identify the time-varying cost parameters. We derive a probabilistic error bound when the cost parameters are identified from finite trajectories. We test our method in numerical and driving simulations. Our algorithm identifies the cost parameters that can reproduce the Nash equilibrium policy and trajectory observations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。