用历史聚合器构建可控制的非马尔可夫决策问题
Constructing Non-Markovian Decision Process via History Aggregator
- 基于范畴论构建马尔可夫与非马尔可夫决策过程的等价框架
- 通过历史聚合器(HAS)精确调控时序决策中的状态依赖结构
- 为强化学习算法评估提供可定制、更严格的非马尔可夫测试环境
在算法决策领域,非马尔可夫动态构成重大障碍,尤其影响强化学习等范式的发展与性能。现有基准在全面评估决策算法处理非马尔可夫动态能力方面存在不足。为此,我们提出一种基于范畴论的通用方法,建立了马尔可夫决策过程(MDP)与非马尔可夫决策过程(NMDP)的范畴,并证明二者之间的等价关系。这一理论基础为理解与应对非马尔可夫动态提供了新视角。我们进一步通过状态历史聚合器(HAS)将非马尔可夫性引入决策问题设置,实现对时序决策中状态依赖结构的精确控制。分析表明,该方法能有效表征多种非马尔可夫动态。该方法使决策算法可在明确构造的非马尔可夫环境中进行更严谨、灵活的评估。
原文摘要 · Abstract (English)
In the domain of algorithmic decision-making, non-Markovian dynamics manifest as a significant impediment, especially for paradigms such as Reinforcement Learning (RL), thereby exerting far-reaching consequences on the advancement and effectiveness of the associated systems. Nevertheless, the existing benchmarks are deficient in comprehensively assessing the capacity of decision algorithms to handle non-Markovian dynamics. To address this deficiency, we have devised a generalized methodology grounded in category theory. Notably, we established the category of Markov Decision Processes (MDP) and the category of non-Markovian Decision Processes (NMDP), and proved the equivalence relationship between them. This theoretical foundation provides a novel perspective for understanding and addressing non-Markovian dynamics. We further introduced non-Markovianity into decision-making problem settings via the History Aggregator for State (HAS). With HAS, we can precisely control the state dependency structure of decision-making problems in the time series. Our analysis demonstrates the effectiveness of our method in representing a broad range of non-Markovian dynamics. This approach facilitates a more rigorous and flexible evaluation of decision algorithms by testing them in problem settings where non-Markovian dynamics are explicitly constructed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。