从观测序列中学习隐藏状态的决策模型,突破传统方法局限。
Toward Learning POMDPs Beyond Full-Rank Actions and State Observability
- 基于谱方法重构部分可观测系统的转移与观测矩阵。
- 在特定秩假设下,可恢复状态划分及对应概率关系。
- 适合需灵活规划的新目标场景,如机械锁控系统建模。
我们关注如何让自主智能体学习并推理具有隐藏状态的系统,例如锁具机制。将此问题建模为学习离散部分可观测马尔可夫决策过程(POMDP)的参数。智能体初始仅知动作和观测空间,未知状态空间、转移或观测模型,这些需从一系列动作与观测中推导。现有谱方法如预测状态表示(PSR)能学习足够预测未来结果的状态表示,但缺乏可用于不同奖励函数的显式转移与观测模型。在对转移与观测矩阵乘积的弱秩假设下,我们证明了可通过张量分解估计相似变换,从而恢复POMDP矩阵。该方法可学习到状态的划分,同一分区内状态对全秩动作具有相同的观测分布。实验表明,学习后的显式观测与转移概率可支持新目标与奖励函数下的新规划。我们进一步通过构造两个观测分布一致但转移动态不同的POMDP,证明仅从序列数据无法超越状态划分学习完整模型。
原文摘要 · Abstract (English)
We are interested in enabling autonomous agents to learn and reason about systems with hidden states, such as locking mechanisms. We cast this problem as learning the parameters of a discrete Partially Observable Markov Decision Process (POMDP). The agent begins with knowledge of the POMDP's actions and observation spaces, but not its state space, transitions, or observation models. These properties must be constructed from a sequence of actions and observations. Spectral approaches to learning models of partially observable domains, such as Predictive State Representations (PSRs), learn representations of state that are sufficient to predict future outcomes. PSR models, however, do not have explicit transition and observation system models that can be used with different reward functions to solve different planning problems. Under a mild set of rankness assumptions on the products of transition and observation matrices, we show how PSRs learn POMDP matrices up to a similarity transform, and this transform may be estimated via tensor decomposition methods. Our method learns observation matrices and transition matrices up to a partition of states, where the states in a single partition have the same observation distributions corresponding to actions whose transition matrices are full-rank. Our experiments suggest that explicit observation and transition likelihoods can be leveraged to generate new plans for different goals and reward functions after the model has been learned. We also show that learning a POMDP beyond a partition of states is impossible from sequential data by constructing two POMDPs that agree on all observation distributions but differ in their transition dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。