arXiv:2502.06051cs.LGcs.AI2025-02被引 8

首次实现反KL正则化离线学习的最优样本复杂度,突破既有理论瓶颈。

Towards a Sharp Analysis of Offline Policy Learning for $f$-Divergence-Regularized Contextual Bandits

  • 基于悲观性分析,首次在单策略集中条件下达成ε⁻¹复杂度
  • 证明单策略集中性对反KL正则化是必要条件,且可最大化利用其曲率优势
  • 适用于追求理论严谨性的强化学习研究者,尤其关注离线算法设计

许多离线强化学习算法依赖于f-散度正则化,但其针对正则化目标的样本复杂度仍缺乏紧致分析,尤其是在具体数据覆盖条件方面。本文研究了实现离线f-散度正则化上下文老虎机中˜Θ(ε⁻¹)样本复杂度所需的精确集中性条件。对于最常用的反KL散度,我们首次在单策略集中性假设下通过新颖的悲观性分析实现了˜O(ε⁻¹)复杂度,优于现有˜O(ε⁻¹)(需全策略集中性)和˜O(ε⁻²)(单策略集中性)的界限。我们还提出了近似匹配的下界,证明乘性依赖单策略集中性是最大化反KL曲率优势所必需的。对于f为强凸的f-散度(反KL不属此类),我们表明即使无需悲观估计或单策略集中性,˜Θ(ε⁻¹)的紧致复杂度也可实现。数值实验验证了理论洞察,并将分析扩展至上下文对抗老虎机。这些结果推动了对f-散度正则化目标的全面理解。

原文摘要 · Abstract (English)

Many offline reinforcement learning algorithms are underpinned by $f$-divergence regularization, but their sample complexity *defined with respect to regularized objectives* still lacks tight analyses, especially in terms of concrete data coverage conditions. In this paper, we study the exact concentrability requirements to achieve the $\tildeΘ(ε^{-1})$ sample complexity for offline $f$-divergence-regularized contextual bandits. For reverse Kullback-Leibler (KL) divergence, arguably the most commonly used one, we achieve an $\tilde{O}(ε^{-1})$ sample complexity under single-policy concentrability for the first time via a novel pessimism-based analysis, surpassing existing $\tilde{O}(ε^{-1})$ bound under all-policy concentrability and $\tilde{O}(ε^{-2})$ bound under single-policy concentrability. We also propose a near-matching lower bound, demonstrating that a multiplicative dependency on single-policy concentrability is necessary to maximally exploit the curvature property of reverse KL. Moreover, for $f$-divergences with strongly convex $f$, to which reverse KL *does not* belong, we show that the sharp sample complexity $\tildeΘ(ε^{-1})$ is achievable even without pessimistic estimation or single-policy concentrability. We further corroborate our theoretical insights with numerical experiments and extend our analysis to contextual dueling bandits. We believe these results take a significant step towards a comprehensive understanding of objectives with $f$-divergence regularization.

强化学习离线学习理论分析散度正则

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。