arXiv:2604.18312cs.LG2026-04ICML被引 4

提出无需知道奖励尺度的自适应规划算法,可高效处理未知奖励环境。

Scale-free adaptive planning for deterministic dynamics & discounted rewards

  • 设计无尺度自适应算法Platypoos,自动适应未知奖励尺度与平滑性
  • 理论证明样本复杂度优于已有方法,且在多种折扣因子下均有效
  • 首次实现对奖励尺度和折扣因子完全未知下的最优性分析

我们研究具有确定性动态和随机奖励、带折扣回报的规划问题。最优价值函数未知,奖励也未受限。本文提出Platypoos——一种简单且无尺度的规划算法,能自适应未知的奖励函数尺度与平滑性。我们为Platypoos提供了样本复杂度分析,其性能优于以往工作,并在广泛折扣因子与奖励尺度下均成立,而算法本身无需预先知晓这些参数。此外,我们建立了匹配的下界,表明该分析在常数意义下已达到最优。

原文摘要 · Abstract (English)

We address the problem of planning in an environment with deterministic dynamics and stochastic rewards with discounted returns. The optimal value function is not known, nor are the rewards bounded. We propose Platypoos, a simple scale-free planning algorithm that adapts to the unknown scale and smoothness of the reward function. We provide a sample complexity analysis for Platypoos that improves upon prior work and holds simultaneously over a broad range of discount factors and reward scales, without the algorithm knowing them. We also establish a matching lower bound showing our analysis is optimal up to constants.

强化学习规划算法无尺度样本复杂度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。