arXiv:2505.16925cs.LG2025-05

用改进的损失函数让强化学习更抗风险,适合高危决策场景。

Risk-Averse Reinforcement Learning with Itakura-Saito Loss

  • 采用伊塔库拉-斋藤散度设计新损失函数,提升数值稳定性。
  • 在多个场景中优于传统方法,尤其在有解析解的测试中表现突出。
  • 适合医疗、金融等需规避高风险的强化学习应用。

风险厌恶型强化学习在诸多高风险领域具有重要应用。与传统强化学习追求期望回报最大化不同,风险厌恶型智能体选择降低风险的策略,有时会牺牲期望收益。这一偏好可通过效用理论建模。本文聚焦指数效用函数情形,可推导出贝尔曼方程,并对现有强化学习算法仅作少量修改即可应用。为此,我们向机器学习社区引入一种基于伊塔库拉-斋藤散度的、数值稳定且数学严谨的损失函数,用于学习状态值函数和动作值函数。我们从理论和实证两方面评估该损失函数,对比了多种已有方法。实验部分涵盖多个场景,包括部分具备解析解的情况,结果表明所提损失函数在性能上显著优于现有替代方案。

原文摘要 · Abstract (English)

Risk-averse reinforcement learning finds application in various high-stakes fields. Unlike classical reinforcement learning, which aims to maximize expected returns, risk-averse agents choose policies that minimize risk, occasionally sacrificing expected value. These preferences can be framed through utility theory. We focus on the specific case of the exponential utility function, where one can derive the Bellman equations and employ various reinforcement learning algorithms with few modifications. To address this, we introduce to the broad machine learning community a numerically stable and mathematically sound loss function based on the Itakura-Saito divergence for learning state-value and action-value functions. We evaluate the Itakura-Saito loss function against established alternatives, both theoretically and empirically. In the experimental section, we explore multiple scenarios, some with known analytical solutions, and show that the considered loss function outperforms the alternatives.

强化学习风险控制损失函数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。