从函数空间视角重新分析PPO,揭示其高效采样机制与理论保证。
Sampling Complexity of TD and PPO in RKHS
- 用核方法分解策略评估与优化,在再生核希尔伯特空间中实现梯度更新。
- 理论证明采样率可达最优的k^{-1/2},在多种模型下统一收敛性保证。
- 适合关注强化学习理论、算法稳定性与样本效率的研究者参考。
我们从函数空间角度重新审视近端策略优化(PPO)。分析将策略评估与改进在再生核希尔伯特空间(RKHS)中解耦:(i) 核化时序差分(TD)评论家仅需单步状态-动作转移样本,即可高效执行RKHS梯度更新;(ii) 通过KL正则化与自然梯度策略步长,对评估的动作价值进行指数化,恢复连续状态-动作空间中的PPO/TRPO式近端更新。我们提供了非渐近、实例自适应的保证,其收敛速率依赖于RKHS熵,统一覆盖表格型、线性、Sobolev、高斯及神经切线核(NTK)情形,并推导出确保随机优化最优k^{-1/2}收敛率的采样规则。实验表明,符合理论的调度策略在常见控制任务(如CartPole、Acrobot)上提升了稳定性和样本效率,且我们的TD评论家相比GAE基线展现出更优吞吐表现。总体而言,本研究在超越有限维假设的前提下,为PPO提供了更坚实的理论基础,阐明了何时使用核TD评论家与RKHS近端更新能实现全局策略提升并保持实际效率。
原文摘要 · Abstract (English)
We revisit Proximal Policy Optimization (PPO) from a function-space perspective. Our analysis decouples policy evaluation and improvement in a reproducing kernel Hilbert space (RKHS): (i) A kernelized temporal-difference (TD) critic performs efficient RKHS-gradient updates using only one-step state-action transition samples; (ii) a KL-regularized, natural-gradient policy step exponentiates the evaluated action-value, recovering a PPO/TRPO-style proximal update in continuous state-action spaces. We provide non-asymptotic, instance-adaptive guarantees whose rates depend on RKHS entropy, unifying tabular, linear, Sobolev, Gaussian, and Neural Tangent Kernel (NTK) regimes, and we derive a sampling rule for the proximal update that ensures the optimal $k^{-1/2}$ convergence rate for stochastic optimization. Empirically, the theory-aligned schedule improves stability and sample efficiency on common control tasks (e.g., CartPole, Acrobot), while our TD-based critic attains favorable throughput versus a GAE baseline. Altogether, our results place PPO on a firmer theoretical footing beyond finite-dimensional assumptions and clarify when RKHS-proximal updates with kernel-TD critics yield global policy improvement with practical efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。