构建首个评估形式化规范强化学习泛化能力的基准测试
SpecRLBench: A Benchmark for Generalization in Specification-Guided Reinforcement Learning

- 设计多难度、跨领域的统一评估框架,涵盖导航与操作任务
- 揭示现有方法在复杂规范和动态环境下的泛化瓶颈
- 适合研究强化学习可扩展性与形式化验证的学者使用
基于形式化规范(如线性时序逻辑,LTL)的强化学习为建模复杂、时序延展的任务提供了系统性框架。尽管近期方法取得进展,其在未见规范与多样化环境中的泛化能力仍不清晰。本文提出SpecRLBench,一个用于评估LTL驱动强化学习方法泛化性能的基准。该基准覆盖导航与操作领域,包含静态与动态环境、多样化的机器人动力学及观测模态,涵盖多个难度层级。通过大量实证评估,我们刻画了现有方法的优劣,并揭示了随着规范与环境复杂度提升所出现的挑战。SpecRLBench为系统性比较与开发更具泛化能力的规范引导强化学习方法提供支持。代码已开源:https://github.com/BU-DEPEND-Lab/SpecRLBench。
原文摘要 · Abstract (English)
Specification-guided reinforcement learning (RL) provides a principled framework for encoding complex, temporally extended tasks using formal specifications such as linear temporal logic (LTL). While recent methods have shown promising results, their ability to generalize across unseen specifications and diverse environments remains insufficiently understood. In this work, we introduce SpecRLBench, a benchmark designed to evaluate the generalization capabilities of LTL-based specification-guided RL methods. The benchmark spans multiple difficulty levels across navigation and manipulation domains, incorporating both static and dynamic environments, diverse robot dynamics, and varied observation modalities. Through extensive empirical evaluation, we characterize the strengths and limitations of existing approaches and reveal the challenges that emerge as specification and environment complexity increase. SpecRLBench provides a structured platform for systematic comparison and supports the development of more generalizable specification-guided RL methods. Code is available at https://github.com/BU-DEPEND-Lab/SpecRLBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。