arXiv:2606.20376cs.LGcs.AI2026-06

CRAX让安全强化学习测试快100倍,支持快速实验与对比。

CRAX: Fast Safe Reinforcement Learning Benchmarking

论文配图:CRAX: Fast Safe Reinforcement Learning Benchmarking
图 1 · 摘自论文原文
  • 基于MJX物理引擎,用向量化和硬件加速实现高速仿真。
  • 比传统CPU基准快约100倍,涵盖6个环境套件、3类任务共9种难度。
  • 适合研究安全RL方法的学者,尤其关注训练策略与安全权衡。

安全是将强化学习(RL)应用于机器人、自动驾驶等现实场景的核心关切。尽管基准测试对推动RL发展至关重要,但现有高保真3D物理安全基准计算效率低,限制了大规模实验与快速原型设计。为此,我们提出CRAX(Constrained RL Accelerated with JAX),基于具有真实3D动态的MuJoCo XLA(MJX)物理引擎,利用向量化操作与硬件加速,相较同类基于CPU的安全基准提速约100倍。该基准包含六个环境套件和三类代理特定任务,每类任务设三个难度级别。评估六种主流安全强化学习方法发现,无单一方法在所有任务上占优,揭示了性能与安全间的权衡。研究还表明,跨难度等级的课程学习与安全迁移可提升在更难设置下的表现。

原文摘要 · Abstract (English)

Safety is a core concern for deploying reinforcement learning (RL) agents in real-world domains such as robotics and autonomous driving. While benchmarks have been central to progress in RL, existing safety benchmarks with high-fidelity 3D physics remain computationally slow, limiting large-scale experimentation and rapid prototyping. To address this gap, we propose CRAX (Constrained RL Accelerated with JAX). Built on top of the MuJoCo XLA (MJX) physics engine with realistic 3D dynamics, CRAX leverages vectorized operations and hardware acceleration, yielding up to ~100x speedups over comparable CPU-based safety benchmarks. The benchmark features six environment suites and three agent-specific tasks, each spanning three difficulty levels. Evaluating six popular safe RL methods shows that no single approach dominates across all tasks, and reveals the trade-offs between performance and safety. We find that curriculum learning across difficulty levels and safety transfer can improve performance over direct training in harder settings.

强化学习安全训练加速仿真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。