仅用轨迹数据即可高效评估策略,突破了传统方法的统计瓶颈。
Trajectory Data Suffices for Statistically Efficient Policy Evaluation in Fixed-Horizon Offline RL with Linear $q^π$-Realizability and Concentrability
- 基于轨迹数据与线性可表示性假设,设计新评估算法
- 实现与最优样本复杂度接近的统计效率
- 适合研究离线强化学习中策略评估的理论工作者
我们研究有限时域离线强化学习中的函数逼近问题,涵盖策略评估与策略优化。以往工作表明,仅在数据覆盖良好(集中性)和所有策略的状态-动作值函数线性可表示(q^π-可表示性)的前提下,两类问题均无法实现统计高效学习(Foster et al., 2021)。近期,Tkachuk 等人(2024)证明:若额外假设数据以轨迹形式提供,则策略优化可实现统计高效。本文在此基础上,首次提出在相同假设下策略评估的统计高效学习算法,并通过更精细分析改进了 Tkachuk 等人所用算法的样本复杂度上界。
原文摘要 · Abstract (English)
We study finite-horizon offline reinforcement learning (RL) with function approximation for both policy evaluation and policy optimization. Prior work established that statistically efficient learning is impossible for either of these problems when the only assumptions are that the data has good coverage (concentrability) and the state-action value function of every policy is linearly realizable ($q^π$-realizability) (Foster et al., 2021). Recently, Tkachuk et al. (2024) gave a statistically efficient learner for policy optimization, if in addition the data is assumed to be given as trajectories. In this work we present a statistically efficient learner for policy evaluation under the same assumptions. Further, we show that the sample complexity of the learner used by Tkachuk et al. (2024) for policy optimization can be improved by a tighter analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。