arXiv:2607.07769cs.LGcs.AI2026-07AAAI

揭示强化学习评估范式缺陷,发现数据越多未必越好

Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms

论文配图:Principled Analysis of Deep Reinforcement Learning Evaluation and Design Paradigms
图 1 · 摘自论文原文
  • 提出强化学习缩放定律理论,分析算法性能与数据规模关系
  • 大规模实验显示现有范式导致错误结论,性能排序非单调
  • 适合关注模型可扩展性与评估可靠性的研究者阅读

从利用深度神经网络逼近状态-动作值函数赢得最复杂游戏,到无需显式规则即可解决难题的算法进步,强化学习在过去十年中推动了显著的科学进展。本文聚焦这一进展的关键要素,分析强化学习中的经典评估与设计范式。我们建立了强化学习缩放定律的理论基础,发现算法的渐近性能与数据规模之间不存在单调关系。通过大规模实验,结果表明在经典范式下的一系列研究得出了错误结论。本分析为深度强化学习的缩放性、容量与复杂性提供了核心洞见。

原文摘要 · Abstract (English)

Starting from the utilization of deep neural networks to approximate the state-action value function that led to winning one of the most challenging games, to algorithmic advancements that allowed solving problems without even explicitly stating the rules of the challenge at hand, reinforcement learning research has been the center of remarkable scientific progress for the past decade. In this paper, we focus on the key ingredients of this research progress and we analyze the canonical evaluation and design paradigms in reinforcement learning. We introduce the theoretical foundations of scaling laws in reinforcement learning and show that the asymptotic performance of reinforcement learning algorithms does not have a monotone relationship between performance rankings and data-regimes. We conduct large-scale experiments and our results demonstrate that a line of reinforcement learning research under the canonical design and evaluation paradigms resulted in incorrect conclusions. Our analysis and results provide a core analysis on scaling, capacity and complexity of deep reinforcement learning.

强化学习评估范式缩放定律

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。