提出一种自适应调节学习速度的新算法,显著加快一般博弈中的无悔学习。
Cautious Optimism: A Meta-Algorithm for Near-Constant Regret in General Games
- 基于乐观学习思想,动态调节学习节奏以加速收敛
- 在多种自对弈场景下实现近最优 $O_T(\log T)$ regret
- 无需知晓对手收益,适合多智能体强化学习应用
我们提出一种名为‘谨慎乐观’(Cautious Optimism)的框架,用于在一般博弈中实现更快的正则化学习。该方法作为乐观策略的变体,以非单调方式自适应地控制学习速度,从而加速无悔学习过程。它接收任意形式的跟随正则化领导者(FTRL)算法作为输入,通过极低计算开销将其转化为加速版无悔学习算法(COFTRL)。关键优势在于保持了不耦合性,即各学习者无需了解其他参与者的效用函数。在多种自对弈场景(混合与匹配正则化项)中,COFTRL 实现了近最优的 $O_T(\log T)$ 无悔值;在对抗性场景下仍保持最优的 $O_T(\sqrt{T})$ 无悔性能。与以往工作(如 Syrgkanis 等 [2015]、Daskalakis 等 [2021])不同,本分析不依赖单调步长,开辟了通用博弈中快速学习的新路径。此外,COFTRL 的具体实例在一般凸博弈中实现了新的最优无悔最小化保证,其对动作空间维度 $d$ 的依赖关系相比先前工作 [Farina 等, 2022a] 实现了指数级改进。
原文摘要 · Abstract (English)
We introduce Cautious Optimism, a framework for substantially faster regularized learning in general games. Cautious Optimism, as a variant of Optimism, adaptively controls the learning pace in a dynamic, non-monotone manner to accelerate no-regret learning dynamics. Cautious Optimism takes as input any instance of Follow-the-Regularized-Leader (FTRL) and outputs an accelerated no-regret learning algorithm (COFTRL) by pacing the underlying FTRL with minimal computational overhead. Importantly, it retains uncoupledness, that is, learners do not need to know other players' utilities. Cautious Optimistic FTRL (COFTRL) achieves near-optimal $O_T(\log T)$ regret in diverse self-play (mixing and matching regularizers) while preserving the optimal $O_T(\sqrt{T})$ regret in adversarial scenarios. In contrast to prior works (e.g., Syrgkanis et al. [2015], Daskalakis et al. [2021]), our analysis does not rely on monotonic step sizes, showcasing a novel route for fast learning in general games. Moreover, instances of COFTRL achieve new state-of-the-art regret minimization guarantees in general convex games, exponentially improving the dependence on the dimension of the action space $d$ over previous works [Farina et al., 2022a].
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。