arXiv:2409.18827cs.LG2024-09被引 5

为强化学习超参优化提供高效可比的基准测试平台。

ARLBench: Flexible and Efficient Benchmarking for Hyperparameter Optimization in Reinforcement Learning

  • 构建跨算法与环境的多样化超参优化任务集
  • 仅需少量计算资源即可生成AutoRL方法性能图谱
  • 适合资源有限的研究者开展强化学习自动化研究

超参数是可靠训练高性能强化学习(RL)智能体的关键因素。然而,开发和评估自动超参数调优方法既耗时又昂贵,导致这些方法通常仅在单一领域或算法上测试,难以比较且限制了对其泛化能力的理解。本文提出ARLBench,一个面向强化学习超参数优化(HPO)的基准测试平台,支持多种HPO方法的高效比较。为促进RL中HPO的研究,即使在计算资源有限的情况下,我们选取了涵盖多种算法与环境组合的代表性子集。该选择使得仅用极少计算量即可生成自动化强化学习(AutoRL)方法的性能轮廓,使更多研究者能参与此领域。基于我们所选任务的广泛大规模超参数景观数据集,ARLBench成为高效、灵活且面向未来的AutoRL研究基础。相关基准与数据集已开源:https://github.com/automl/arlbench。

原文摘要 · Abstract (English)

Hyperparameters are a critical factor in reliably training well-performing reinforcement learning (RL) agents. Unfortunately, developing and evaluating automated approaches for tuning such hyperparameters is both costly and time-consuming. As a result, such approaches are often only evaluated on a single domain or algorithm, making comparisons difficult and limiting insights into their generalizability. We propose ARLBench, a benchmark for hyperparameter optimization (HPO) in RL that allows comparisons of diverse HPO approaches while being highly efficient in evaluation. To enable research into HPO in RL, even in settings with low compute resources, we select a representative subset of HPO tasks spanning a variety of algorithm and environment combinations. This selection allows for generating a performance profile of an automated RL (AutoRL) method using only a fraction of the compute previously necessary, enabling a broader range of researchers to work on HPO in RL. With the extensive and large-scale dataset on hyperparameter landscapes that our selection is based on, ARLBench is an efficient, flexible, and future-oriented foundation for research on AutoRL. Both the benchmark and the dataset are available at https://github.com/automl/arlbench.

强化学习超参优化AutoRL基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。