揭示了策略学习中重置分布对样本复杂度的决定性影响。
The Sample Complexity of Policy Learning with Mu-Resets
- 通过分析重置分布的覆盖能力,确定样本复杂度的关键因素。
- 在有界全策略集中性下,样本复杂度呈指数级增长(exp(Ω(H)))。
- 适用于研究强化学习理论、样本效率及环境设计的研究者。
我们研究了基于策略的强化学习在 Kakade 和 Langford [KL02] 提出的 μ-重置交互协议下的表现。该协议允许学习者从给定的探索性重置分布 μ 中采样轨迹,而不仅限于初始分布。本文解决了 [KLS25] 提出的关于策略可实现性对样本复杂度影响的问题。关键发现是:与时长远近相关的复杂度取决于重置分布的覆盖假设。在有界全策略集中性条件下,我们建立了 exp(Ω(H)) 的样本复杂度下界;而在有界前向集中性条件下,其依赖关系被精确刻画为 exp(Θ(√H))。
原文摘要 · Abstract (English)
We study policy-based reinforcement learning under the $μ$-resets interaction protocol of Kakade and Langford [KL02]. This interaction protocol enables the learner to sample trajectories from a given exploratory reset distribution $μ$, in addition to the starting distribution. We resolve the question raised by [KLS25] on the role of policy realizability for the sample complexity of this problem. Critically, the dependence on horizon $H$ is governed by the notion of coverage assumed of the reset distribution. Under bounded all-policy concentrability, we show a $\exp(Ω(H))$ sample complexity lower bound; with bounded pushforward concentrability, we show the dependence on horizon is tightly characterized as $\exp(Θ(\sqrt H))$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。