arXiv:2603.03191stat.MLcs.LG2026-03

用信念空间度量改进离线POMDP的评估方法,提升样本效率。

A Covering Framework for Offline POMDPs Learning using Belief Space Metric

  • 基于信念空间度量构建覆盖分析框架,替代传统历史覆盖
  • 在双采样贝尔曼误差和记忆依赖价值函数中实现更紧的误差界
  • 适合研究离线强化学习与部分可观测决策问题的学者

在部分可观测马尔可夫决策过程(POMDP)的离线策略评估(OPE)中,智能体需从过往观测推断隐藏状态,这加剧了现有方法中的有限时域与记忆诅咒问题。本文提出一种新颖的覆盖分析框架,利用信念空间(隐状态分布)的内在度量结构,放松传统覆盖假设。通过假设价值相关函数在信念空间中满足Lipschitz连续性,我们推导出可缓解时域与记忆长度指数增长的误差界。该统一分析方法适用于广泛的OPE算法,得到以信念空间度量表达的误差界与覆盖要求,而非原始历史覆盖。案例研究显示:双采样贝尔曼误差最小化算法与基于记忆的未来依赖价值函数(FDVF),在基于信念空间度量的覆盖定义下均获得更紧的误差界,显著提升样本效率。

原文摘要 · Abstract (English)

In off policy evaluation (OPE) for partially observable Markov decision processes (POMDPs), an agent must infer hidden states from past observations, which exacerbates both the curse of horizon and the curse of memory in existing OPE methods. This paper introduces a novel covering analysis framework that exploits the intrinsic metric structure of the belief space (distributions over latent states) to relax traditional coverage assumptions. By assuming value relevant functions are Lipschitz continuous in the belief space, we derive error bounds that mitigate exponential blow ups in horizon and memory length. Our unified analysis technique applies to a broad class of OPE algorithms, yielding concrete error bounds and coverage requirements expressed in terms of belief space metrics rather than raw history coverage. We illustrate the improved sample efficiency of this framework via case studies: the double sampling Bellman error minimization algorithm, and the memory based future dependent value functions (FDVF). In both cases, our coverage definition based on the belief space metric yields tighter bounds.

POMDP离线评估信念空间强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。