arXiv:2507.00257cs.LGcs.AI2025-07被引 4

构建真实世界强化学习基准测试套件,推动算法落地应用。

Gym4ReaL: A Suite for Benchmarking Real-World Reinforcement Learning

  • 设计包含多种现实挑战的模拟环境,覆盖大状态空间与部分可观测性。
  • 标准RL算法在真实场景中表现优于规则基线,验证其竞争力。
  • 适合关注算法真实部署的开发者与研究者使用。

近年来,强化学习(RL)在各类模拟环境中取得了显著进展,甚至达到超人类表现。但当研究转向真实世界应用时,面临状态-动作空间庞大、非平稳性、部分可观测性等挑战。现有基准多聚焦理想化、完全可观测且平稳的环境,未能显式涵盖真实世界的复杂性。本文提出 exttt{Gym4ReaL},一个全面的现实环境套件,旨在支持可在真实场景中运行的强化学习算法开发与评估。该套件包含多样任务,使算法暴露于实际挑战中。实验结果表明,在这些设置下,标准RL算法在性能上仍可超越规则基线,激励新方法的发展以充分应对真实任务的复杂性。

原文摘要 · Abstract (English)

In recent years, \emph{Reinforcement Learning} (RL) has made remarkable progress, achieving superhuman performance in a wide range of simulated environments. As research moves toward deploying RL in real-world applications, the field faces a new set of challenges inherent to real-world settings, such as large state-action spaces, non-stationarity, and partial observability. Despite their importance, these challenges are often underexplored in current benchmarks, which tend to focus on idealized, fully observable, and stationary environments, often neglecting to incorporate real-world complexities explicitly. In this paper, we introduce \texttt{Gym4ReaL}, a comprehensive suite of realistic environments designed to support the development and evaluation of RL algorithms that can operate in real-world scenarios. The suite includes a diverse set of tasks that expose algorithms to a variety of practical challenges. Our experimental results show that, in these settings, standard RL algorithms confirm their competitiveness against rule-based benchmarks, motivating the development of new methods to fully exploit the potential of RL to tackle the complexities of real-world tasks.

强化学习真实世界基准测试算法评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。