构建九个工业级优化任务基准,推动安全强化学习落地现实决策
SafeOR-Gym: A Benchmark Suite for Safe Reinforcement Learning Algorithms on Practical Operations Research Problems
- 设计九个含混合离散连续动作的复杂优化环境,模拟真实工业约束
- 多算法测试显示当前安全强化学习在部分任务上表现不佳
- 专为能源、制造等高风险领域设计,适合研究实际约束下的安全决策
现有安全强化学习基准主要聚焦于机器人与控制任务,难以覆盖涉及结构化约束、混合整数决策和工业复杂性的高风险领域。这一缺口制约了安全强化学习在能源系统、制造和供应链等关键场景的应用推进。为此,我们提出 SafeOR-Gym,一个包含九个运筹学(OR)环境的基准套件,专为复杂约束下的安全强化学习设计。每个环境均刻画真实的规划、调度或控制问题,具有基于成本的约束违规机制、规划时域以及混合离散-连续动作空间。该套件无缝集成 OmniSafe 提供的约束马尔可夫决策过程(CMDP)接口。我们在这些环境中评估多个前沿安全强化学习算法,结果呈现显著性能差异:部分任务可解,但另一些暴露出现有方法的根本局限。SafeOR-Gym 提供了一个兼具挑战性与实用性的测试平台,旨在推动安全强化学习在真实世界决策中的未来发展。
原文摘要 · Abstract (English)
Most existing safe reinforcement learning (RL) benchmarks focus on robotics and control tasks, offering limited relevance to high-stakes domains that involve structured constraints, mixed-integer decisions, and industrial complexity. This gap hinders the advancement and deployment of safe RL in critical areas such as energy systems, manufacturing, and supply chains. To address this limitation, we present SafeOR-Gym, a benchmark suite of nine operations research (OR) environments tailored for safe RL under complex constraints. Each environment captures a realistic planning, scheduling, or control problems characterized by cost-based constraint violations, planning horizons, and hybrid discrete-continuous action spaces. The suite integrates seamlessly with the Constrained Markov Decision Process (CMDP) interface provided by OmniSafe. We evaluate several state-of-the-art safe RL algorithms across these environments, revealing a wide range of performance: while some tasks are tractable, others expose fundamental limitations in current approaches. SafeORGym provides a challenging and practical testbed that aims to catalyze future research in safe RL for real-world decision-making problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。