arXiv:2508.10340cs.AI2025-08被引 1

动态分配信任域阈值,让多智能体强化学习更稳定高效

Multi-Agent Trust Region Policy Optimisation: A Joint Constraint Approach

  • 按智能体优化信任域阈值分配,结合全局约束与贪心策略
  • 在多个基准上提升超22.5%最终奖励,收敛更快且方差更低
  • 适合复杂异构多智能体系统,如协作博弈与竞争环境

多智能体强化学习(MARL)需要智能体间协调稳定的策略更新。异构智能体信任域策略优化(HATRPO)使用相对熵(KL)约束每个智能体的策略更新以稳定训练。然而,在异构设置下,给所有智能体设定相同KL阈值会导致更新缓慢且陷入局部最优。为此,我们提出两种阈值分配方法:基于KKT的HATRPO-W,在全局KL约束下优化阈值分配;以及基于改进率/发散比优先级的贪心算法HATRPO-G。通过将序列策略优化与约束阈值调度结合,本方法在异构场景中实现更灵活有效的学习。实验表明,所提方法显著提升HATRPO性能,在多个MARL基准上实现更快收敛与更高最终奖励,其中两者均使最终性能提升超过22.5%。值得注意的是,HATRPO-W还表现出更稳定的训练动态,方差更低。

原文摘要 · Abstract (English)

Multi-agent reinforcement learning (MARL) requires coordinated and stable policy updates among interacting agents. Heterogeneous-Agent Trust Region Policy Optimization (HATRPO) enforces per-agent trust region constraints using Kullback-Leibler (KL) divergence to stabilize training. However, assigning each agent the same KL threshold can lead to slow and locally optimal updates, especially in heterogeneous settings. To address this limitation, we propose two approaches for allocating the KL divergence threshold across agents: HATRPO-W, a Karush-Kuhn-Tucker-based (KKT-based) method that optimizes threshold assignment under global KL constraints, and HATRPO-G, a greedy algorithm that prioritizes agents based on improvement-to-divergence ratio. By connecting sequential policy optimization with constrained threshold scheduling, our approach enables more flexible and effective learning in heterogeneous-agent settings. Experimental results demonstrate that our methods significantly boost the performance of HATRPO, achieving faster convergence and higher final rewards across diverse MARL benchmarks. Specifically, HATRPO-W and HATRPO-G achieve comparable improvements in final performance, each exceeding 22.5%. Notably, HATRPO-W also demonstrates more stable learning dynamics, as reflected by its lower variance.

多智能体强化学习信任域策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。