arXiv:2410.15156cs.AIcs.MA2024-10被引 2

多智能体博弈中用KL代价优化策略,可独立计算最优解。

Simulation-Based Optimistic Policy Iteration For Multi-Agent MDPs with Kullback-Leibler Control Cost

  • 基于贝尔曼分布设计独立更新的策略改进机制
  • 同步与异步版本均渐近收敛到最优价值与策略
  • 适用于需控制代价的多智能体协同场景

本文提出一种面向多智能体马尔可夫决策过程(MDPs)的基于代理的乐观策略迭代(OPI)方案,用于学习静态最优随机策略。各智能体在执行控制时产生Kullback-Leibler(KL)发散代价,并承担联合状态的额外成本。该方案包含贪婪策略改进步骤与m步时间差分(TD)策略评估步骤。利用即时成本的可分结构,证明策略改进遵循依赖当前值函数估计与无控制转移概率的玻尔兹曼分布,使各智能体可独立计算改进后的联合策略。我们证明了同步(全状态空间评估)与异步(均匀采样的子状态集)版本的OPI在有限策略评估展开下,均渐近收敛至最优值函数与最优联合策略。在带有KL控制代价的猎鹿-野兔博弈变体上的仿真结果验证了该方案在最小化总成本回报方面的有效性。

原文摘要 · Abstract (English)

This paper proposes an agent-based optimistic policy iteration (OPI) scheme for learning stationary optimal stochastic policies in multi-agent Markov Decision Processes (MDPs), in which agents incur a Kullback-Leibler (KL) divergence cost for their control efforts and an additional cost for the joint state. The proposed scheme consists of a greedy policy improvement step followed by an m-step temporal difference (TD) policy evaluation step. We use the separable structure of the instantaneous cost to show that the policy improvement step follows a Boltzmann distribution that depends on the current value function estimate and the uncontrolled transition probabilities. This allows agents to compute the improved joint policy independently. We show that both the synchronous (entire state space evaluation) and asynchronous (a uniformly sampled set of substates) versions of the OPI scheme with finite policy evaluation rollout converge to the optimal value function and an optimal joint policy asymptotically. Simulation results on a multi-agent MDP with KL control cost variant of the Stag-Hare game validates our scheme's performance in terms of minimizing the cost return.

多智能体策略迭代KL代价博弈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。