用KL正则化实现无需人为设定悲观惩罚的多智能体离线学习。
Pessimism-Free Offline Learning in General-Sum Games via KL Regularization
- 通过KL正则化替代人工悲观惩罚,稳定学习过程。
- 理论证明可加速达到统计收敛率$\widetilde{O}(1/n)$。
- 适合研究多智能体博弈中无偏离线学习的学者参考。
在一般和博弈的离线多智能体强化学习中,日志数据分布与目标均衡策略之间存在分布偏移问题。传统方法依赖手动设定悲观惩罚,我们证明仅使用KL正则化即可稳定学习并实现均衡恢复。提出通用锚定纳什均衡(GANE),以$\widetilde{O}(1/n)$的加速统计速率恢复正则化纳什均衡。为保证计算可行性,设计通用锚定镜面下降(GAMD)算法,迭代收敛至粗相关均衡,速率为$\widetilde{O}(1/\sqrt{n}+1/T)$。结果表明,KL正则化可作为无悲观性离线学习的独立机制,在多玩家一般和博弈中实现等效或更快的收敛速度。
原文摘要 · Abstract (English)
Offline multi-agent reinforcement learning in general-sum settings is challenged by the distribution shift between logged datasets and target equilibrium policies. While standard methods rely on manual pessimistic penalties, we demonstrate that KL regularization suffices to stabilize learning and achieve equilibrium recovery. We propose General-sum Anchored Nash Equilibrium (GANE), which recovers regularized Nash equilibria at an accelerated statistical rate of $\widetilde{O}(1/n)$. For computational tractability, we develop General-sum Anchored Mirror Descent (GAMD), an iterative algorithm converging to a Coarse Correlated Equilibrium at the standard rate of $\widetilde{O}(1/\sqrt{n}+1/T)$. These results establish KL regularization as a standalone mechanism for pessimism-free offline learning that achieves equivalent or accelerated rates in multi-player general-sum games.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。