提出无需悲观估计的离线学习方法,实现更快收敛速率。
Beyond Pessimism: Offline Learning in KL-regularized Games
- 基于KL正则化策略的平滑最优响应和纳什均衡稳定性设计新算法
- 首次实现无悲观性保证的离线学习,样本复杂度达$ ilde{ m O}(1/n)$
- 适用于需高效稳定训练的博弈型强化学习场景
我们研究在KL正则化双人零和博弈中的离线学习,其中策略通过KL正则化相对于固定参考策略进行优化。以往工作依赖悲观价值估计处理分布偏移,仅能获得$ ilde{ m O}(1/ oot ext{ }n)$的统计率。本文提出一种新的无悲观性算法与分析框架,利用KL正则化最优响应的光滑性及由偏对称性诱导的纳什均衡稳定性,首次实现针对KL正则化博弈的无悲观性离线学习保证,达到$ ilde{ m O}(1/n)$的快速样本复杂度。此外,我们提出一种高效的自对弈策略优化算法,以迭代的KL正则化策略更新替代精确均衡计算,并证明其最终迭代仍保持相同的无悲观性统计保证,误差可控。
原文摘要 · Abstract (English)
We study offline learning in KL-regularized two-player zero-sum games, where policies are optimized with respect to a fixed reference policy through KL regularization. Prior work relies on pessimistic value estimation to handle distribution shift, yielding only $\widetilde{\mathcal{O}}(1/\sqrt n)$ statistical rates. We develop a new pessimism-free algorithm and analytical framework for KL-regularized games, built on the smoothness of KL-regularized best responses and a stability property of the Nash equilibrium induced by skew symmetry. This yields, to our knowledge, the first pessimism-free offline learning guarantee for KL-regularized games, with a fast $\widetilde{\mathcal{O}}(1/n)$ sample complexity bound. We further propose an efficient self-play policy optimization algorithm that replaces exact equilibrium computation with iterative KL-regularized policy updates, and prove that its last iterate preserves the same pessimism-free statistical guarantee up to a controlled optimization error.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。