提出评估强化学习超参敏感性的新方法,揭示部分性能提升实为调参依赖。
A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning
- 设计可量化算法对超参变化响应的实验方法
- 发现多种PPO归一化变体敏感度差异显著
- 适合关注模型稳定性和调参效率的研究者
现代强化学习算法的性能高度依赖不断增长的超参数数量。超参微小变动常导致性能剧烈波动,且不同环境需迥异的超参设置才能达到文献报告的最优表现。当前缺乏可扩展且广泛接受的方法来刻画这些复杂交互关系。本文提出一种新的经验性方法,用于研究、比较和量化特定环境下算法性能对超参调优的敏感性。通过该方法评估了多个常用PPO归一化变体的敏感性,结果表明,部分算法性能提升可能实质上源于对超参数调优的更强依赖。
原文摘要 · Abstract (English)
The performance of modern reinforcement learning algorithms critically relies on tuning ever-increasing numbers of hyperparameters. Often, small changes in a hyperparameter can lead to drastic changes in performance, and different environments require very different hyperparameter settings to achieve state-of-the-art performance reported in the literature. We currently lack a scalable and widely accepted approach to characterizing these complex interactions. This work proposes a new empirical methodology for studying, comparing, and quantifying the sensitivity of an algorithm's performance to hyperparameter tuning for a given set of environments. We then demonstrate the utility of this methodology by assessing the hyperparameter sensitivity of several commonly used normalization variants of PPO. The results suggest that several algorithmic performance improvements may, in fact, be a result of an increased reliance on hyperparameter tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。