arXiv:2603.00374cs.AIcs.MA2026-03

在离线多智能体博弈中,通过保守策略选择找到更接近真实均衡的解。

Conservative Equilibrium Discovery in Offline Game-Theoretic Multiagent Reinforcement Learning

  • 基于策略空间响应原理由在线算法扩展,引入不确定性量化与保守性目标
  • 在多个博弈数据集上实现比现有方法更低的遗憾值(regret)
  • 适合研究离线多智能体强化学习或博弈求解的学者使用

离线学习通过仅依赖固定状态-动作轨迹数据集,将数据效率推至极致。本文研究在混合动机多智能体场景下的离线博弈求解问题,核心是于候选均衡中进行选择。由于数据集仅能反映部分博弈动态,通常无法验证所提解是否为真实均衡。因此,我们基于可用信息,评估各候选解产生低遗憾(即接近均衡)的相对概率。具体地,将在线博弈求解方法政策空间响应原器(PSRO)扩展为可量化博弈动态不确定性的版本,并修改强化学习目标,使解向更可能在真实博弈中表现低遗憾的方向倾斜。此外,提出一种专为离线设定设计的新元策略求解器,引导PSRO中的策略探索。该方法结合了离线强化学习中的保守性原则,命名为COffeE-PSRO。实验表明,该方法能提取出比当前最优离线方法更低遗憾的解,并揭示算法组件、经验博弈保真度与整体性能之间的关系。

原文摘要 · Abstract (English)

Offline learning of strategies takes data efficiency to its extreme by restricting algorithms to a fixed dataset of state-action trajectories. We consider the problem in a mixed-motive multiagent setting, where the goal is to solve a game under the offline learning constraint. We first frame this problem in terms of selecting among candidate equilibria. Since datasets may inform only a small fraction of game dynamics, it is generally infeasible in offline game-solving to even verify a proposed solution is a true equilibrium. Therefore, we consider the relative probability of low regret (i.e., closeness to equilibrium) across candidates based on the information available. Specifically, we extend Policy Space Response Oracles (PSRO), an online game-solving approach, by quantifying game dynamics uncertainty and modifying the RL objective to skew towards solutions more likely to have low regret in the true game. We further propose a novel meta-strategy solver, tailored for the offline setting, to guide strategy exploration in PSRO. Our incorporation of Conservatism principles from Offline reinforcement learning approaches for strategy Exploration gives our approach its name: COffeE-PSRO. Experiments demonstrate COffeE-PSRO's ability to extract lower-regret solutions than state-of-the-art offline approaches and reveal relationships between algorithmic components empirical game fidelity, and overall performance.

多智能体离线学习博弈论策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。