arXiv:2409.12799stat.MLcs.LG2024-09中稿 · Statistical Scienc…被引 15

损失函数决定强化学习效率,选对损失能大幅提升算法表现。

The Central Role of the Loss Function in Reinforcement Learning

  • 用二元交叉熵损失可实现最优策略成本的首阶边界,比平方损失更高效。
  • 分布式RL采用最大似然损失时,达到次阶边界,性能优于传统方法。
  • 适合研究决策算法优化或想提升强化学习模型表现的读者。

本文系统阐述了损失函数在数据驱动决策中的核心作用,综述其在代价敏感分类(CSC)和强化学习(RL)中的影响。我们证明,在多种设置下,使用二元交叉熵损失的算法能实现与最优策略成本相关的首阶边界,显著优于常用的平方损失。此外,采用最大似然损失的分布式算法能达到与策略方差相关的次阶边界,性能更优。这一结果特别验证了分布式强化学习的优势。本文旨在为分析不同损失函数下的决策算法提供指导,并激励读者探索更优损失函数以改进各类决策系统。

原文摘要 · Abstract (English)

This paper illustrates the central role of loss functions in data-driven decision making, providing a comprehensive survey on their influence in cost-sensitive classification (CSC) and reinforcement learning (RL). We demonstrate how different regression loss functions affect the sample efficiency and adaptivity of value-based decision making algorithms. Across multiple settings, we prove that algorithms using the binary cross-entropy loss achieve first-order bounds scaling with the optimal policy's cost and are much more efficient than the commonly used squared loss. Moreover, we prove that distributional algorithms using the maximum likelihood loss achieve second-order bounds scaling with the policy variance and are even sharper than first-order bounds. This in particular proves the benefits of distributional RL. We hope that this paper serves as a guide analyzing decision making algorithms with varying loss functions, and can inspire the reader to seek out better loss functions to improve any decision making algorithm.

强化学习损失函数分布式RL算法优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。