arXiv:2603.28681stat.MLcs.LG2026-03被引 1

提出新方法,让离线强化学习在复杂策略下仍能高效收敛。

Functional Natural Policy Gradients

  • 用交叉拟合去偏机制提升离线策略学习稳定性
  • 在策略复杂度超Donsker时仍保证√N的后悔界
  • 适用于高复杂度环境下的策略优化研究

我们提出一种用于从离线数据中进行策略学习的交叉拟合去偏装置。该学习原则的关键结果是,即使策略类的复杂度超过Donsker条件,只要误差乘积的干扰项为O(N⁻¹/²),仍可实现√N的后悔界。该后悔界分解为由策略类复杂度决定的插值策略误差因子和由环境动态复杂度决定的环境干扰因子,明确揭示了二者之间的权衡关系。

原文摘要 · Abstract (English)

We propose a cross-fitted debiasing device for policy learning from offline data. A key consequence of the resulting learning principle is $\sqrt N$ regret even for policy classes with complexity greater than Donsker, provided a product-of-errors nuisance remainder is $O(N^{-1/2})$. The regret bound factors into a plug-in policy error factor governed by policy-class complexity and an environment nuisance factor governed by the complexity of the environment dynamics, making explicit how one may be traded against the other.

强化学习离线学习策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。