arXiv:2604.01024cs.LG2026-04被引 1

提出高效学习有限窗口策略的方法,让部分可观测问题更易解。

Model-Based Learning of Near-Optimal Finite-Window Policies in POMDPs

  • 用历史窗口构建马尔可夫决策过程,简化部分可观测问题
  • 仅需单条轨迹即可保证模型估计的样本效率
  • 适合需要稳定策略的机器人控制等实际场景

我们研究在表格型部分可观测马尔可夫决策过程(POMDP)中基于模型的学习,采用有限动作-观测窗口来近似无界历史依赖。这将原始问题转化为一个关于历史序列的有限状态马尔可夫决策过程(称为超状态MDP)。一旦获得该超状态MDP的模型,即可使用标准MDP算法求解最优策略。由于轨迹来自原POMDP的交互,采样过程与目标模型之间存在不匹配,导致模型估计困难。本文提出一种针对表格型POMDP的模型估计方法,并分析其样本复杂度。分析利用滤波稳定性与弱相关随机变量的集中不等式之间的联系,从而得到从单一轨迹中估计超状态MDP模型的紧致样本复杂度上界。结合值迭代,可得到近似最优的有限窗口策略。

原文摘要 · Abstract (English)

We study model-based learning of finite-window policies in tabular partially observable Markov decision processes (POMDPs). A common approach to learning under partial observability is to approximate unbounded history dependencies using finite action-observation windows. This induces a finite-state Markov decision process (MDP) over histories, referred to as the superstate MDP. Once a model of this superstate MDP is available, standard MDP algorithms can be used to compute optimal policies, motivating the need for sample-efficient model estimation. Estimating the superstate MDP model is challenging because trajectories are generated by interaction with the original POMDP, creating a mismatch between the sampling process and target model. We propose a model estimation procedure for tabular POMDPs and analyze its sample complexity. Our analysis exploits a connection between filter stability and concentration inequalities for weakly dependent random variables. As a result, we obtain tight sample complexity guarantees for estimating the superstate MDP model from a single trajectory. Combined with value iteration, this yields approximately optimal finite-window policies for the POMDP.

POMDP策略学习样本效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。