arXiv:2501.16600cs.GTcs.LG2025-01被引 1

改进强化学习中的噪声稳定性,让算法更稳定收敛。

On the Power of Perturbation under Sampling in Solving Extensive-Form Games

  • 引入扰动的FTRL算法,通过KL与反KL两种方式优化学习过程
  • 反KL版本在莱德克扑克中表现更优,提升最后迭代收敛性
  • 适合研究博弈学习、鲁棒强化学习的学者参考

我们研究在采样环境下,扰动对求解不完美信息扩展形式博弈中跟随正则化领导者(FTRL)算法的影响。尽管乐观算法在完整反馈下有效,但在采样噪声下常不稳定。扰动提供了一种稳定学习并实现最后迭代收敛的替代方案。本文提出统一的扰动FTRL框架,包含两种变体:基于标准KL散度的PFTRL-KL和基于反KL散度的PFTRL-RKL,后者具备无偏估计与条件零方差特性。在基准博弈中,PFTRL-KL表现相当或更优;而在结构更不对称的莱德克扑克中,PFTRL-RKL持续领先。第二组实验隔离了条件零方差效应,验证其能有效降低方差,提升最后迭代性能。

原文摘要 · Abstract (English)

We investigate how perturbation does and does not improve the Follow-the-Regularized-Leader (FTRL) algorithm in solving imperfect-information extensive-form games under sampling, where payoffs are estimated from sampled trajectories. While optimistic algorithms are effective under full feedback, they often become unstable in the presence of sampling noise. Payoff perturbation offers a promising alternative for stabilizing learning and achieving \textit{last-iterate convergence}. We present a unified framework for \textit{Perturbed FTRL} algorithms and study two variants: PFTRL-KL (standard KL divergence) and PFTRL-RKL (Reverse KL divergence), the latter featuring an estimator with both unbiasedness and conditional zero variance. While PFTRL-KL generally achieves equivalent or better performance across benchmark games, PFTRL-RKL consistently outperforms it in Leduc poker, whose structure is more asymmetric than the other games in a sense. Given the modest advantage of PFTRL-RKL, we design the second experiment to isolate the effect of conditional zero variance, showing that the variance-reduction property of RKL improve last-iterate performance.

博弈学习强化学习收敛性扰动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。