arXiv:2409.15703eess.SYcs.LG2024-09被引 14

提出统一框架,用可递归更新的代理状态替代信念状态,提升部分可观测强化学习效果。

Agent-state based policies in POMDPs: Beyond belief-state MDPs

  • 以可递归更新的代理状态代替传统信念状态,简化决策过程。
  • 通过近似信息状态改进Q-learning与演员-评论家算法在部分可观测环境中的表现。
  • 适用于模型未知场景,为强化学习中的部分可观测问题提供新思路。

传统POMDP求解方法将问题转化为基于信念状态的完全可观测MDP,但要求已知系统动态,难以应用于模型未知的学习场景。本文提出一种统一视角,将多种方法视为代理维持局部可递归更新的代理状态并据此决策的模型。文中梳理了不同类别的代理状态策略:设计者构造的最优非平稳策略、策略搜索得到的局部最优平稳策略,以及近似信息状态实现的近似最优平稳策略。进一步说明,近似信息状态的思想已被用于改进Q-learning和演员-评论家算法,在无模型条件下实现更优的POMDP学习性能。

原文摘要 · Abstract (English)

The traditional approach to POMDPs is to convert them into fully observed MDPs by considering a belief state as an information state. However, a belief-state based approach requires perfect knowledge of the system dynamics and is therefore not applicable in the learning setting where the system model is unknown. Various approaches to circumvent this limitation have been proposed in the literature. We present a unified treatment of some of these approaches by viewing them as models where the agent maintains a local recursively updateable agent state and chooses actions based on the agent state. We highlight the different classes of agent-state based policies and the various approaches that have been proposed in the literature to find good policies within each class. These include the designer's approach to find optimal non-stationary agent-state based policies, policy search approaches to find a locally optimal stationary agent-state based policies, and the approximate information state to find approximately optimal stationary agent-state based policies. We then present how ideas from the approximate information state approach have been used to improve Q-learning and actor-critic algorithms for learning in POMDPs.

强化学习部分可观测策略优化智能体状态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。