arXiv:2504.18917cs.GTcs.LG2025-04被引 1

用元学习提升自对弈中的后悔最小化,全局优化策略更高效。

Meta-Learning in Self-Play Regret Minimization

  • 元学习融合多博弈状态信息,实现全局策略协同。
  • 在正常形式博弈与河牌扑克子游戏中性能超越现有算法。
  • 适合大规模博弈中需快速适应新场景的玩家策略优化。

后悔最小化是在线优化的通用方法,在近似双人零和博弈纳什均衡的算法中起关键作用。现有研究主要关注孤立求解单个博弈,但实际中玩家常面临一系列相似但不同的博弈分布,如股票市场中相关资产交易或大型博弈的子游戏策略优化。近期已有离线元学习用于加速单边均衡发现。本文在此基础上,将框架扩展至更具挑战性的自对弈设置,这正是大多数大规模领域最优均衡近似算法的基础。我们的方法在选择策略时,独特地整合所有决策状态的信息,促进全局通信而非传统局部后悔分解。在正常形式博弈与河牌扑克子游戏上的实验表明,元学习算法显著优于其他最先进的后悔最小化算法。

原文摘要 · Abstract (English)

Regret minimization is a general approach to online optimization which plays a crucial role in many algorithms for approximating Nash equilibria in two-player zero-sum games. The literature mainly focuses on solving individual games in isolation. However, in practice, players often encounter a distribution of similar but distinct games. For example, when trading correlated assets on the stock market, or when refining the strategy in subgames of a much larger game. Recently, offline meta-learning was used to accelerate one-sided equilibrium finding on such distributions. We build upon this, extending the framework to the more challenging self-play setting, which is the basis for most state-of-the-art equilibrium approximation algorithms for domains at scale. When selecting the strategy, our method uniquely integrates information across all decision states, promoting global communication as opposed to the traditional local regret decomposition. Empirical evaluation on normal-form games and river poker subgames shows our meta-learned algorithms considerably outperform other state-of-the-art regret minimization algorithms.

元学习自对弈博弈优化后悔最小化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。