arXiv:2410.23498cs.LGcs.AI2024-10NeurIPS被引 7

用核方法提升平均奖励强化学习的预测能力,实现无遗憾学习。

Kernel-Based Function Approximation for Average Reward Reinforcement Learning: An Optimist No-Regret Algorithm

  • 基于核岭回归构建价值函数近似模型,提升表达能力。
  • 在无限时域平均奖励设置下,证明算法达到无遗憾性能。
  • 提出通用核预测置信区间,适用于多种强化学习问题。

利用核岭回归预测期望价值函数的强化学习方法具有强大的表示能力,且适合进行理论分析。本文研究了在无限时域平均奖励(即未折现)设置下的核基函数逼近方法。我们提出一种乐观算法,其思想类似于带参考项的老虎机算法。在核基建模假设下,建立了该算法的新型无遗憾性能保证。此外,我们推导出适用于各类强化学习问题的核基价值函数预测的新置信区间。

原文摘要 · Abstract (English)

Reinforcement learning utilizing kernel ridge regression to predict the expected value function represents a powerful method with great representational capacity. This setting is a highly versatile framework amenable to analytical results. We consider kernel-based function approximation for RL in the infinite horizon average reward setting, also referred to as the undiscounted setting. We propose an optimistic algorithm, similar to acquisition function based algorithms in the special case of bandits. We establish novel no-regret performance guarantees for our algorithm, under kernel-based modelling assumptions. Additionally, we derive a novel confidence interval for the kernel-based prediction of the expected value function, applicable across various RL problems.

强化学习核方法无遗憾

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。