arXiv:2602.06599cs.MAcs.AI2026-02中稿 · the 25th Internati…

提出联合经验最优响应,大幅降低多智能体强化学习的样本开销。

Sample-Efficient Policy Space Response Oracles with Joint Experience Best Response

  • 收集一次联合轨迹,同时计算所有智能体的最优响应
  • 探索增强版JBR在基准环境上实现最高精度与效率平衡
  • 适合大规模多智能体系统中需要高效策略迭代的研究者

多智能体强化学习(MARL)为可扩展的游戏理论分析提供替代方案,但面临非平稳性及需维护多样策略种群以捕捉非传递性交互的问题。策略空间响应算子(PSRO)通过迭代扩展受限博弈并引入近似最优响应(BR)来解决上述问题,但在多智能体或模拟器昂贵的场景下,单智能体最优响应训练成本过高。本文提出联合经验最优响应(JBR),作为PSRO的即插即用改进:仅需在当前元策略配置下收集一次联合轨迹,并复用该数据集同时计算所有智能体的最优响应,从而摊薄环境交互成本,显著提升最优响应计算的样本效率。由于JBR将最优响应计算转为离线强化学习问题,我们提出三种缓解分布偏移偏差的方法:(i) 保守型JBR采用安全策略改进,(ii) 探索增强型JBR对数据收集过程扰动,具备理论保障,(iii) 混合型最优响应交替使用JBR与周期性独立最优响应更新。在多个基准多智能体环境中,探索增强型JBR达到最佳精度-效率权衡,混合型方法在仅消耗极小样本的情况下接近原始PSRO性能。总体而言,JBR使PSRO在大规模战略学习中更具实用性,同时保持均衡鲁棒性。

原文摘要 · Abstract (English)

Multi-agent reinforcement learning (MARL) offers a scalable alternative to exact game-theoretic analysis but suffers from non-stationarity and the need to maintain diverse populations of strategies that capture non-transitive interactions. Policy Space Response Oracles (PSRO) address these issues by iteratively expanding a restricted game with approximate best responses (BRs), yet per-agent BR training makes it prohibitively expensive in many-agent or simulator-expensive settings. We introduce Joint Experience Best Response (JBR), a drop-in modification to PSRO that collects trajectories once under the current meta-strategy profile and reuses this joint dataset to compute BRs for all agents simultaneously. This amortizes environment interaction and improves the sample efficiency of best-response computation. Because JBR converts BR computation into an offline RL problem, we propose three remedies for distribution-shift bias: (i) Conservative JBR with safe policy improvement, (ii) Exploration-Augmented JBR that perturbs data collection and admits theoretical guarantees, and (iii) Hybrid BR that interleaves JBR with periodic independent BR updates. Across benchmark multi-agent environments, Exploration-Augmented JBR achieves the best accuracy-efficiency trade-off, while Hybrid BR attains near-PSRO performance at a fraction of the sample cost. Overall, JBR makes PSRO substantially more practical for large-scale strategic learning while preserving equilibrium robustness.

多智能体强化学习样本效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。