arXiv:2605.07218cs.LGstat.ML2026-05

用平滑核方法改进连续空间强化学习,理论更优。

Improved Model-based Reinforcement Learning with Smooth Kernels

论文配图:Improved Model-based Reinforcement Learning with Smooth Kernels
图 1 · 摘自论文原文
  • 结合伯恩斯坦探索奖励与核平滑,提升学习效率
  • 在有限时域下实现更优的累积误差上界
  • 适合关注理论性能的强化学习研究者

针对连续状态-动作空间场景,传统强化学习理论多基于低秩马尔可夫决策过程(MDP),虽保证样本高效但结构假设过强。核平滑模型基方法提供新范式,利用MDP的平滑性,采用非参数核平滑估计转移动态。本文提出一种新的核平滑模型基在线强化学习方法,适用于满足利普希茨连续性假设的有限时域设置。通过在核平滑框架中引入伯恩斯坦风格探索奖励,该方法实现了优于现有最优结果的后悔上界,其对时域的依赖关系更优。理论突破源于对伯恩斯坦奖励与核平滑协同效应的精细分析,其中提出的新型紧致伯恩斯坦型鞅浓度不等式或具独立研究价值。

原文摘要 · Abstract (English)

For continuous state-action space scenarios, classical reinforcement learning (RL) theory predominantly focuses on low-rank Markov decision processes (MDPs), which provide sample-efficient guarantees at the expense of restrictive structural assumptions. Kernel smoothing model-based approaches offer a promising alternative paradigm that instead leverages the smoothness of the MDP and employs non-parametric kernel smoothing estimates of transition dynamics. This paper proposes a new kernel-smoothing model-based approach for online reinforcement learning in finite-horizon settings under Lipschitz continuity assumptions on the MDP. By incorporating a Bernstein-style exploration bonus into the kernel smoothing framework, our method achieves a regret bound which improves upon the state-of-the-art regret bound in its dependence on the horizon. The theoretical advancement relies on a delicate analysis of the synergy between Bernstein-style bonuses and kernel smoothing, where a new tight Bernstein-type concentration inequality for martingales may be of independent interest.

强化学习核方法理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。