arXiv:2502.07715cs.LG2025-02中稿 · AISTATS 2025被引 1

提出高效奖励无关的核方法强化学习,显著降低样本需求。

Near-Optimal Sample Complexity in Reward-Free Kernel-Based Reinforcement Learning

  • 基于核函数构建新置信区间,简化算法设计
  • 在非线性环境中实现近优策略仅需较少采样
  • 适合研究样本效率与复杂函数逼近的学者

强化学习正面临日益复杂的结构挑战。尽管表格和线性模型已得到充分研究,但具有强表示能力和理论可分析性的核函数模型在非线性函数逼近中的研究近年兴起。本文在奖励无关强化学习框架下,探讨核函数强化学习的统计效率问题:需要多少样本才能设计出近优策略?现有工作受限于对核函数类别的严格假设。我们首先在生成模型假设下分析该问题,随后放宽假设,代价是样本复杂度增加一个因子H(episode长度)。我们采用广义核函数类与更简洁的算法,推导出适用于本强化学习场景的新核岭回归置信区间,可能具更广泛适用性。并通过仿真验证了理论结果。

原文摘要 · Abstract (English)

Reinforcement Learning (RL) problems are being considered under increasingly more complex structures. While tabular and linear models have been thoroughly explored, the analytical study of RL under nonlinear function approximation, especially kernel-based models, has recently gained traction for their strong representational capacity and theoretical tractability. In this context, we examine the question of statistical efficiency in kernel-based RL within the reward-free RL framework, specifically asking: how many samples are required to design a near-optimal policy? Existing work addresses this question under restrictive assumptions about the class of kernel functions. We first explore this question by assuming a generative model, then relax this assumption at the cost of increasing the sample complexity by a factor of H, the length of the episode. We tackle this fundamental problem using a broad class of kernels and a simpler algorithm compared to prior work. Our approach derives new confidence intervals for kernel ridge regression, specific to our RL setting, which may be of broader applicability. We further validate our theoretical findings through simulations.

强化学习核方法样本效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。