arXiv:2501.09646cs.AI2025-01被引 2

首个专为非平稳马尔可夫决策过程设计的开源仿真与评测工具

NS-Gym: Open-Source Simulation Environments and Benchmarks for Non-Stationary Markov Decision Processes

  • 将环境参数变化与智能体决策分离,实现模块化动态适应
  • 提供标准化接口和基准任务,支持算法在非平稳环境中的可复现评估
  • 适合研究自适应决策、强化学习鲁棒性的学者使用

在众多现实应用中,智能体需在受外部因素影响而不断变化的环境中做出序列决策。传统决策模型通常假设环境动态是平稳的,难以应对此类非平稳环境。非平稳马尔可夫决策过程(NS-MDPs)为此类问题提供了建模框架。然而,缺乏标准基准和仿真工具严重制约了该领域的系统性研究进展。本文提出NS-Gym,首个专为NS-MDPs设计的开源仿真工具包,集成于流行的Gymnasium框架。在NS-Gym中,环境参数演化与智能体决策模块分离,支持灵活的动态适应。我们梳理了该领域前期工作,构建了涵盖关键特征与类型的问题集,并首次建立了一套标准化接口与基准任务,以实现非平稳条件下算法评估的一致性与可复现性。我们还使用NS-Gym对六种已有算法进行了基准测试。愿景是使研究人员能够评估其决策算法在非平稳环境中的适应性与鲁棒性。

原文摘要 · Abstract (English)

In many real-world applications, agents must make sequential decisions in environments where conditions are subject to change due to various exogenous factors. These non-stationary environments pose significant challenges to traditional decision-making models, which typically assume stationary dynamics. Non-stationary Markov decision processes (NS-MDPs) offer a framework to model and solve decision problems under such changing conditions. However, the lack of standardized benchmarks and simulation tools has hindered systematic evaluation and advance in this field. We present NS-Gym, the first simulation toolkit designed explicitly for NS-MDPs, integrated within the popular Gymnasium framework. In NS-Gym, we segregate the evolution of the environmental parameters that characterize non-stationarity from the agent's decision-making module, allowing for modular and flexible adaptations to dynamic environments. We review prior work in this domain and present a toolkit encapsulating key problem characteristics and types in NS-MDPs. This toolkit is the first effort to develop a set of standardized interfaces and benchmark problems to enable consistent and reproducible evaluation of algorithms under non-stationary conditions. We also benchmark six algorithmic approaches from prior work on NS-MDPs using NS-Gym. Our vision is that NS-Gym will enable researchers to assess the adaptability and robustness of their decision-making algorithms to non-stationary conditions.

强化学习非平稳环境仿真工具决策系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。