将强化学习中的评估方法扩展到分布级,提升离线策略评估的准确性。
A Principled Path to Fitted Distributional Evaluation
- 基于拟合技巧构建分布级评估框架,统一设计新方法。
- 在线性二次调节器与Atari游戏上验证,性能优于现有方法。
- 提供理论支持,适用于非表格型复杂环境,适合研究者参考。
在强化学习中,分布级离线策略评估(OPE)旨在利用不同策略收集的离线数据估计目标策略的回报分布。本文将广泛使用的拟合Q评估方法(用于期望值评估)扩展至分布级设置,提出拟合分布评估(FDE)。尽管已有少数相关方法,但尚无统一的FDE方法设计框架。为此,本文提出一系列指导原则,用于构建理论严谨的FDE方法。基于这些原则,我们开发了几种新的FDE方法,并提供了收敛性分析,同时为现有方法提供了理论解释,即使在非表格型环境中也成立。大量实验,包括在线性二次调节器和Atari游戏上的模拟,证明了FDE方法的优越性能。
原文摘要 · Abstract (English)
In reinforcement learning, distributional off-policy evaluation (OPE) focuses on estimating the return distribution of a target policy using offline data collected under a different policy. This work focuses on extending the widely used fitted Q-evaluation -- developed for expectation-based reinforcement learning -- to the distributional OPE setting. We refer to this extension as fitted distributional evaluation (FDE). While only a few related approaches exist, there remains no unified framework for designing FDE methods. To fill this gap, we present a set of guiding principles for constructing theoretically grounded FDE methods. Building on these principles, we develop several new FDE methods with convergence analysis and provide theoretical justification for existing methods, even in non-tabular environments. Extensive experiments, including simulations on linear quadratic regulators and Atari games, demonstrate the superior performance of the FDE methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。