用稀疏加性模型提升强化学习离线评估的精度与可解释性
Sparse Additive Off-Policy Evaluation for Reinforcement Learning with Potentially Limited Number of Trajectories
- 基于稀疏加性结构建模Q函数,支持非线性且可解释的策略评估
- 在轨迹数或时长足够大时仍能保证高概率误差界,缓解维度灾难
- 提出分组稀疏特征筛选法,高效识别关键影响变量,适合小样本场景
我们提出一种灵活、非线性且可解释的无限时域强化学习离线评估新框架。为应对大规模状态空间并支持透明决策,采用具有稀疏加性结构的非线性函数类建模Q函数。推导出目标策略价值函数估计的高概率有限样本误差界,其依赖于环境维度d的对数项,有效缓解维度灾难。与多数现有理论假设大量轨迹不同,本分析证明:当轨迹数量或时间跨度足够大时,即可实现准确的价值估计。此外,提出基于分组稀疏性的特征筛选方法,在高概率下识别出包含所有相关协变量的缩减特征集。数值实验验证了该方法的有效性。
原文摘要 · Abstract (English)
We develop a new framework for flexible, nonlinear, and interpretable off-policy evaluation for infinite-horizon reinforcement learning. To handle large state spaces and support transparent decision-making, we model the Q-function using a nonlinear function class with a sparse additive structure. We derive high-probability finite-sample error bounds for estimating the value function of a target policy and show that the bounds depend only logarithmically on the ambient dimension $d$, thereby alleviating the curse of dimensionality. In contrast to most existing theory for off-policy evaluation, which typically assumes access to many trajectories, our analysis guarantees accurate value estimation when either the number of trajectories or the time horizon is sufficiently large. In addition, we propose a group-sparsity-based feature screening procedure that identifies, with high probability, a reduced feature set containing all relevant covariates. Numerical experiments demonstrate the effectiveness of the proposed approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。