arXiv:2603.10895cs.LG2026-03中稿 · article to appear …

非遍历性奖励让传统RL目标失效,论文提出改进方法。

Ergodicity in reinforcement learning

  • 用实例说明非遍历奖励导致期望值失效
  • 发现单条轨迹长期表现与平均值可能不一致
  • 提供针对非遍历动态的优化方案,适合关注个体性能的研究者

强化学习通常以最大化智能体在轨迹上累积奖励的期望值为目标。然而,若生成奖励的过程是非遍历性的,该期望值(即基于给定策略的无限多轨迹平均)对单条无限长轨迹的平均表现无意义。因此,若关心部署时个体智能体的表现,期望值并非合适优化目标。本文通过一个直观示例探讨非遍历奖励过程对强化学习智能体的影响,将非遍历奖励过程与更常见的马尔可夫链遍历性概念关联,并总结现有解决非遍历奖励动态下个体长期性能优化的方法。

原文摘要 · Abstract (English)

In reinforcement learning, we typically aim to optimize the expected value of the sum of rewards an agent collects over a trajectory. However, if the process generating these rewards is non-ergodic, the expected value, i.e., the average over infinitely many trajectories with a given policy, is uninformative for the average over a single, but infinitely long trajectory. Thus, if we care about how the individual agent performs during deployment, the expected value is not a good optimization objective. In this paper, we discuss the impact of non-ergodic reward processes on reinforcement learning agents through an instructive example, relate the notion of ergodic reward processes to more widely used notions of ergodic Markov chains, and present existing solutions that optimize long-term performance of individual trajectories under non-ergodic reward dynamics.

强化学习非遍历性奖励设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。