构建统一模块化基准,评估强化学习在多种干扰下的鲁棒性。
Robust Gymnasium: A Unified Modular Benchmark for Robust Reinforcement Learning
- 设计可模块化扩展的基准框架,支持状态、奖励、动作和环境多维度干扰。
- 涵盖60+种任务环境,覆盖控制、机器人、安全强化学习与多智能体场景。
- 揭示现有算法在鲁棒性上的明显短板,适合研究鲁棒强化学习的学者使用。
由于固有的不确定性及仿真到现实的差距,鲁棒强化学习(RL)旨在提升智能体在复杂多变的交互序列中的抗扰能力。尽管已有大量强化学习基准,但缺乏针对鲁棒强化学习的标准化评测体系。当前的鲁棒算法多针对特定类型不确定性,且在孤立环境中评估。本文提出 Robust-Gymnasium,一个统一的模块化基准,支持对智能体观测状态与奖励、动作以及环境本身施加多样化的干扰。该平台提供超过六十种多样化任务环境,涵盖控制与机器人、安全强化学习及多智能体强化学习领域,为社区提供开源、易用的评估工具,以检验现有方法并推动鲁棒强化学习算法的发展。此外,我们在该框架下对主流标准与鲁棒强化学习算法进行了基准测试,揭示了各类算法存在的显著缺陷,并提供了新的洞察。
原文摘要 · Abstract (English)
Driven by inherent uncertainty and the sim-to-real gap, robust reinforcement learning (RL) seeks to improve resilience against the complexity and variability in agent-environment sequential interactions. Despite the existence of a large number of RL benchmarks, there is a lack of standardized benchmarks for robust RL. Current robust RL policies often focus on a specific type of uncertainty and are evaluated in distinct, one-off environments. In this work, we introduce Robust-Gymnasium, a unified modular benchmark designed for robust RL that supports a wide variety of disruptions across all key RL components-agents' observed state and reward, agents' actions, and the environment. Offering over sixty diverse task environments spanning control and robotics, safe RL, and multi-agent RL, it provides an open-source and user-friendly tool for the community to assess current methods and foster the development of robust RL algorithms. In addition, we benchmark existing standard and robust RL algorithms within this framework, uncovering significant deficiencies in each and offering new insights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。