arXiv:2601.19030cs.LGcs.AI2026-01被引 2

提出新覆盖度量,统一线性离策略评估的误差分析框架。

A Unifying View of Coverage in Linear Off-Policy Evaluation

  • 从工具变量视角重构算法误差,引入特征-动态覆盖新概念。
  • 在仅目标函数线性可表示条件下,给出严格有限样本误差界。
  • 统一了经典设定与弱假设下的覆盖定义,适合理论研究者参考。

离策略评估(OPE)是强化学习的基础任务。在线性OPE的经典设定中,有限样本保证常表现为:评估误差 ≤ poly(C^π, d, 1/n, log(1/δ)),其中d为特征维度,C^π为刻画访问特征在数据分布张量空间中覆盖程度的覆盖参数。尽管在更强假设(如贝尔曼完备性)下多个算法的保证已清晰,但在仅目标值函数线性可表示的最简设定下,统计率的紧致刻画仍不明确,既有覆盖度量存在缺陷且与文献中标准定义脱节。本文对典型算法LSTDQ进行新颖的有限样本分析,受工具变量视角启发,提出依赖于新覆盖参数——特征-动态覆盖的误差界。该参数可解释为特征演化动态系统中的线性覆盖。在附加贝尔曼完备性等假设下,其能自然恢复已有特定场景的覆盖定义,最终实现线性离策略评估中覆盖概念的统一理解。

原文摘要 · Abstract (English)

Off-policy evaluation (OPE) is a fundamental task in reinforcement learning (RL). In the classic setting of linear OPE, finite-sample guarantees often take the form $$ \textrm{Evaluation error} \le \textrm{poly}(C^π, d, 1/n,\log(1/δ)), $$ where $d$ is the dimension of the features and $C^π$ is a coverage parameter that characterizes the degree to which the visited features lie in the span of the data distribution. While such guarantees are well-understood for several popular algorithms under stronger assumptions (e.g. Bellman completeness), the understanding is lacking and fragmented in the minimal setting where only the target value function is linearly realizable in the features. Despite recent interest in tight characterizations of the statistical rate in this setting, the right notion of coverage remains unclear, and candidate definitions from prior analyses have undesirable properties and are starkly disconnected from more standard definitions in the literature. We provide a novel finite-sample analysis of a canonical algorithm for this setting, LSTDQ. Inspired by an instrumental-variable view, we develop error bounds that depend on a novel coverage parameter, the feature-dynamics coverage, which can be interpreted as linear coverage in an induced dynamical system for feature evolution. With further assumptions -- such as Bellman-completeness -- our definition successfully recovers the coverage parameters specialized to those settings, finally yielding a unified understanding for coverage in linear OPE.

强化学习离策略评估理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。