arXiv:2508.01329cs.LGcs.AI2025-08被引 1

深度强化学习的瓶颈是优化不足,而非探索不够。

Is Exploration or Optimization the Problem for Deep Reinforcement Learning?

  • 提出实用的次优性评估方法,量化优化能力限制。
  • 实验显示最佳经验性能比实际策略高2-3倍。
  • 适合关注训练效率与算法改进的RL研究者。

在深度强化学习时代,将收集的经验压缩进深度模型以供未来利用和采样变得愈发复杂。许多研究指出,在状态与动作分布不断变化的情况下训练深度学习策略,会导致次优性能甚至模型崩溃。这引发了一个核心问题:即便社区开发出更优的探索算法或奖励目标,这些改进是否因优化困难而被忽视?本文提出一种新的实用次优性评估方法,用于判断深度强化学习算法的优化局限性。在多个环境和强化学习算法上的实验表明,最优生成经验的性能比策略实际学习到的性能高出2-3倍。这一显著差距说明,深度强化学习方法仅利用了其生成的良好经验的一半左右。

原文摘要 · Abstract (English)

In the era of deep reinforcement learning, making progress is more complex, as the collected experience must be compressed into a deep model for future exploitation and sampling. Many papers have shown that training a deep learning policy under the changing state and action distribution leads to sub-optimal performance, or even collapse. This naturally leads to the concern that even if the community creates improved exploration algorithms or reward objectives, will those improvements fall on the \textit{deaf ears} of optimization difficulties. This work proposes a new \textit{practical} sub-optimality estimator to determine optimization limitations of deep reinforcement learning algorithms. Through experiments across environments and RL algorithms, it is shown that the difference between the best experience generated is 2-3$\times$ better than the policies' learned performance. This large difference indicates that deep RL methods only exploit half of the good experience they generate.

强化学习优化瓶颈经验利用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。