arXiv:2410.20092cs.LGcs.AI2024-10ICLR被引 217

构建首个系统化评估离线目标条件强化学习的基准测试平台

OGBench: Benchmarking Offline Goal-Conditioned RL

  • 设计8类环境与85个数据集,覆盖多种复杂场景
  • 揭示现有算法在长时序推理等能力上的显著差异
  • 适合强化学习研究者和算法开发者参考

离线目标条件强化学习(Offline GCRL)是强化学习中的重要问题,因其能无需奖励信号、以无监督方式从无标签数据中学习多样化行为与表征。然而,当前缺乏系统评估该领域算法的标准化基准。本文提出OGBench,一个高质量的离线目标条件强化学习基准平台,包含8类环境、85个数据集,以及6种代表性算法的参考实现。这些环境与数据集设计具有挑战性且贴近真实场景,可直接检验算法在行为拼接、长时序推理、高维输入处理及随机性应对等方面的能力。实验表明,尽管主流算法在已有基准上表现相近,但在OGBench上却显现出明显优劣差异,为新算法研发提供了坚实基础。

原文摘要 · Abstract (English)

Offline goal-conditioned reinforcement learning (GCRL) is a major problem in reinforcement learning (RL) because it provides a simple, unsupervised, and domain-agnostic way to acquire diverse behaviors and representations from unlabeled data without rewards. Despite the importance of this setting, we lack a standard benchmark that can systematically evaluate the capabilities of offline GCRL algorithms. In this work, we propose OGBench, a new, high-quality benchmark for algorithms research in offline goal-conditioned RL. OGBench consists of 8 types of environments, 85 datasets, and reference implementations of 6 representative offline GCRL algorithms. We have designed these challenging and realistic environments and datasets to directly probe different capabilities of algorithms, such as stitching, long-horizon reasoning, and the ability to handle high-dimensional inputs and stochasticity. While representative algorithms may rank similarly on prior benchmarks, our experiments reveal stark strengths and weaknesses in these different capabilities, providing a strong foundation for building new algorithms. Project page: https://seohong.me/projects/ogbench

强化学习离线学习基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。