arXiv:2410.08870cs.LG2024-10被引 9

不同版本的Hopper环境让强化学习算法表现差异巨大,需科学选择基准。

Can we hop in general? A discussion of benchmark selection and design using the Hopper environment

  • 通过对比Hopper环境变体,揭示基准选择对算法评估的影响。
  • 实验表明不同环境设定下算法性能排名可能完全反转。
  • 呼吁建立统一标准来解释和讨论基准环境的选择依据。

强化学习(RL)研究中依赖实证基准测试是普遍做法,但基准选择常基于直觉,如‘有腿机器人’或‘视觉观测’,缺乏系统性讨论。本文以Hopper环境为例,展示不同变体如何显著改变对算法性能的评判结果。当前研究并未形成对各类Hopper环境代表性的共识,甚至彼此之间也不具可比性。实验结果表明,深度强化学习文献中基准选择既缺乏普遍解释,也缺少用于说明选择合理性的语言体系。本文最后讨论了科学评估基准所需的标准,并建议采取措施推动该领域的对话与规范建立。

原文摘要 · Abstract (English)

Empirical, benchmark-driven testing is a fundamental paradigm in the current RL community. While using off-the-shelf benchmarks in reinforcement learning (RL) research is a common practice, this choice is rarely discussed. Benchmark choices are often done based on intuitive ideas like "legged robots" or "visual observations". In this paper, we argue that benchmarking in RL needs to be treated as a scientific discipline itself. To illustrate our point, we present a case study on different variants of the Hopper environment to show that the selection of standard benchmarking suites can drastically change how we judge performance of algorithms. The field does not have a cohesive notion of what the different Hopper environments are representative - they do not even seem to be representative of each other. Our experimental results suggests a larger issue in the deep RL literature: benchmark choices are neither commonly justified, nor does there exist a language that could be used to justify the selection of certain environments. This paper concludes with a discussion of the requirements for proper discussion and evaluations of benchmarks and recommends steps to start a dialogue towards this goal.

强化学习基准测试环境设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。