为非平稳策略的强化学习提供可信赖的置信区间估计方法
Model-based Bootstrap of Controlled Markov Chains
- 基于模型的自助法重构转移核,适配历史依赖策略
- 在河泳任务中,百分位置信区间覆盖率接近名义水平
- 适合小样本、短序列下的离线策略评估与最优策略恢复
我们提出并分析了一种针对有限可控马尔可夫链(CMC)转移核的基于模型的自助法,适用于可能非平稳或依赖历史的控制策略,该设定在行为策略未知的离线强化学习中自然出现。我们建立了单条长轨迹和分幕式离线强化学习两种情形下自助法转移估计器的分布一致性。关键技术工具包括:针对访问频次的新自助大数定律,以及对自助转移增量的鞅中心极限定理新应用。通过验证贝尔曼算子的Hadamard可微性,利用delta方法将自助分布一致性扩展至下游目标,包括离线策略评估(OPE)和最优策略恢复(OPR),得到价值函数和$Q$-函数的渐近有效置信区间。在河泳(RiverSwim)问题上的实验表明,所提自助置信区间(尤其是百分位法)优于分幕自助法和插件式中心极限定理置信区间,在小样本和短轨迹条件下仍能保持接近名义覆盖水平(50%、90%、95%),而基线方法则严重失准。
原文摘要 · Abstract (English)
We propose and analyze a model-based bootstrap for transition kernels in finite controlled Markov chains (CMCs) with possibly nonstationary or history-dependent control policies, a setting that arises naturally in offline reinforcement learning (RL) when the behavior policy generating the data is unknown. We establish distributional consistency of the bootstrap transition estimator in both a single long-chain regime and the episodic offline RL regime. The key technical tools are a novel bootstrap law of large numbers (LLN) for the visitation counts and a novel use of the martingale central limit theorem (CLT) for the bootstrap transition increments. We extend bootstrap distributional consistency to the downstream targets of offline policy evaluation (OPE) and optimal policy recovery (OPR) via the delta method by verifying Hadamard differentiability of the Bellman operators, yielding asymptotically valid confidence intervals for value and $Q$-functions. Experiments on the RiverSwim problem show that the proposed bootstrap confidence intervals (CIs), especially the percentile CIs, outperform the episodic bootstrap and plug-in CLT CIs, and are often close to nominal ($50\%$, $90\%$, $95\%$) coverage, while the baselines are poorly calibrated at small sample sizes and short episode lengths.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。