RL研究需从比拼性能转向理解学习机制,避免盲目追求榜单突破。
The Formalism-Implementation Gap in Reinforcement Learning Research
- 主张减少对性能榜单的依赖,转而深入研究RL的学习机理。
- 指出当前基准测试与数学理论间存在脱节,影响技术可迁移性。
- 以ALE为例,说明经典环境仍可用于深化对RL的理解。
过去十年,强化学习(RL)因在部分任务上达到'超人类水平'而受到广泛关注和应用,这促使学术界更重视展示智能体性能的研究,而忽视了对其学习动态的理解。这种以性能为导向的研究容易在学术基准上过拟合,降低其实际可用性,并阻碍技术向新问题的迁移。同时,这类研究也无形中贬低了那些不追求性能前沿但致力于提升对算法理解的工作。本文提出两点主张:(i) RL研究应超越单纯展示智能体能力,更多关注科学探索与机制理解;(ii) 需更精确地界定基准测试与其背后数学形式之间的映射关系。我们以流行的雅典娜学习环境(ALE; Bellemare et al., 2013)为例,说明尽管该环境常被认为已'饱和',但仍可有效用于深化对强化学习的理解,并推动其在现实世界中的应用。
原文摘要 · Abstract (English)
The last decade has seen an upswing in interest and adoption of reinforcement learning (RL) techniques, in large part due to its demonstrated capabilities at performing certain tasks at "super-human levels". This has incentivized the community to prioritize research that demonstrates RL agent performance, often at the expense of research aimed at understanding their learning dynamics. Performance-focused research runs the risk of overfitting on academic benchmarks -- thereby rendering them less useful -- which can make it difficult to transfer proposed techniques to novel problems. Further, it implicitly diminishes work that does not push the performance-frontier, but aims at improving our understanding of these techniques. This paper argues two points: (i) RL research should stop focusing solely on demonstrating agent capabilities, and focus more on advancing the science and understanding of reinforcement learning; and (ii) we need to be more precise on how our benchmarks map to the underlying mathematical formalisms. We use the popular Arcade Learning Environment (ALE; Bellemare et al., 2013) as an example of a benchmark that, despite being increasingly considered "saturated", can be effectively used for developing this understanding, and facilitating the deployment of RL techniques in impactful real-world problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。