arXiv:2509.13653cs.GTcs.LG2025-09

改进奖励变换方法,让博弈求解更快收敛到稳定策略。

Efficient Last-Iterate Convergence in Regret Minimization via Adaptive Reward Transformation

  • 动态调整参数,自动平衡探索与利用。
  • 实现线性收敛,比现有方法快数倍。
  • 适用于各类博弈模型,尤其适合复杂场景。

后悔最小化是寻找正常形式博弈(NFG)和广义形式博弈(EFG)纳什均衡的强大方法,但通常仅能保证平均策略的收敛。计算平均策略需要大量计算资源或引入额外误差,限制了实际应用。奖励变换(RT)框架通过奖励函数正则化实现了逐轮策略收敛,但其性能高度依赖人工调参,常偏离理论收敛条件,导致收敛缓慢、振荡或陷入局部最优。受先前工作启发,我们提出自适应技术,使RT后悔匹配(RTRM)、RT反事实后悔最小化(RTCFR)及其变体在解决NFG和EFG时更具一致性与鲁棒性。所提方法动态调节参数,在改善后悔累积的同时提升渐近最后一轮策略收敛性,实现线性收敛。实验表明,该方法显著加速收敛,优于当前最先进算法。

原文摘要 · Abstract (English)

Regret minimization is a powerful method for finding Nash equilibria in Normal-Form Games (NFGs) and Extensive-Form Games (EFGs), but it typically guarantees convergence only for the average strategy. However, computing the average strategy requires significant computational resources or introduces additional errors, limiting its practical applicability. The Reward Transformation (RT) framework was introduced to regret minimization to achieve last-iterate convergence through reward function regularization. However, it faces practical challenges: its performance is highly sensitive to manually tuned parameters, which often deviate from theoretical convergence conditions, leading to slow convergence, oscillations, or stagnation in local optima. Inspired by previous work, we propose an adaptive technique to address these issues, ensuring better consistency between theoretical guarantees and practical performance for RT Regret Matching (RTRM), RT Counterfactual Regret Minimization (RTCFR), and their variants in solving NFGs and EFGs more effectively. Our adaptive methods dynamically adjust parameters, balancing exploration and exploitation while improving regret accumulation, ultimately enhancing asymptotic last-iterate convergence and achieving linear convergence. Experimental results demonstrate that our methods significantly accelerate convergence, outperforming state-of-the-art algorithms.

博弈求解后悔最小化自适应优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。