arXiv:2503.01134cs.LGcs.AI2025-03ICLR被引 8

证明了在部分可观测环境下,历史依赖策略的离线评估难以准确实现。

Statistical Tractability of Off-policy Evaluation of History-dependent Policies in POMDPs

  • 通过信息论分析,揭示了模型无关方法在历史依赖策略上的理论局限性。
  • 提出简单模型基算法可突破限制,实现可证明的准确评估。
  • 适合研究强化学习评估理论与复杂环境建模的学者参考。

我们研究了在具有大观测空间的部分可观测马尔可夫决策过程(POMDPs)中,离线策略评估(OPE)这一强化学习核心问题。近期工作(Uehara et al., 2023a;Zhang & Jiang, 2024)提出了无模型框架,并识别出关键覆盖假设(信念覆盖与结果覆盖),使得对无记忆策略的准确OPE可在多项式样本复杂度下实现,但对依赖完整可观测历史的更一般目标策略仍属开放问题。本文证明,在多种设定下,对历史依赖策略的无模型OPE存在信息论上的困难,这些设定由行为策略(无记忆或历史依赖)和状态揭示特性(单步或多步揭示)的额外假设决定。我们进一步表明,通过一个自然的模型基算法可克服部分困难——其分析虽简单却长期未被文献关注——从而在POMDPs中证明了无模型与模型基OPE之间的可证明差异。

原文摘要 · Abstract (English)

We investigate off-policy evaluation (OPE), a central and fundamental problem in reinforcement learning (RL), in the challenging setting of Partially Observable Markov Decision Processes (POMDPs) with large observation spaces. Recent works of Uehara et al. (2023a); Zhang & Jiang (2024) developed a model-free framework and identified important coverage assumptions (called belief and outcome coverage) that enable accurate OPE of memoryless policies with polynomial sample complexities, but handling more general target policies that depend on the entire observable history remained an open problem. In this work, we prove information-theoretic hardness for model-free OPE of history-dependent policies in several settings, characterized by additional assumptions imposed on the behavior policy (memoryless vs. history-dependent) and/or the state-revealing property of the POMDP (single-step vs. multi-step revealing). We further show that some hardness can be circumvented by a natural model-based algorithm -- whose analysis has surprisingly eluded the literature despite the algorithm's simplicity -- demonstrating provable separation between model-free and model-based OPE in POMDPs.

强化学习离线评估部分可观测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。