arXiv:2605.11289cs.LGmath.OC2026-05被引 1

解决平均奖励强化学习中分布估计的数学病态问题。

Quotient-Categorical Representations for Bellman-Compatible Average-Reward Distributional Reinforcement Learning

论文配图:Quotient-Categorical Representations for Bellman-Compatible Average-Reward Distributional Reinforcement Learning
图 1 · 摘自论文原文
  • 用商空间与分类参数化处理偏差的平移对称性
  • 证明了投影算子在坐标Cramér度量下非扩张且存在不动点
  • 适用于在线估增益、马尔可夫采样等实际场景

平均奖励强化学习需同时估计增益和偏差,但偏差仅在加性常数意义下定义,导致实数轴上的直接分布式类比病态。本文提出一种商空间形式化,将状态相关的偏差律视为在共同平移下等价,并引入满足该对称性的分类参数化。在此商-分类空间上,定义了一个投影平均奖励分布算子,证明其在坐标Cramér度量下为非扩张且存在不动点。研究了均值场映射为该算子异步松弛的采样递归,在理想中心化奖励设定下,单状态TD更新在独立同分布与马尔可夫采样下均实现几乎必然收敛及有限迭代残差界。当增益未知时,通过在线增益估计器扩展递归,证明了耦合方案的非扩张性与马尔可夫收敛性。最后表明,同步精确更新在商律层面与增益无关,揭示了理想商分布与实际固定网格分类表示间的结构性差异。

原文摘要 · Abstract (English)

Average-reward reinforcement learning requires estimating the gain and the bias, which is defined only up to an additive constant. This makes direct distributional analogues ill-posed on the real line. We introduce a quotient-space formulation in which state-indexed bias laws are identified up to a common translation, together with a categorical parameterization that respects this symmetry. On this quotient-categorical space, we define a projected average-reward distributional operator and show that it is well-defined, non-expansive in a coordinate Cramér metric, and admits fixed points. We then study sampled recursions whose mean-field maps are asynchronous relaxations of this operator. In an idealized centered-reward setting, a one-state temporal-difference update enjoys almost sure convergence together with finite-iteration residual bounds under both i.i.d. and Markovian sampling. When the gain is unknown, we augment the recursion with an online gain estimator, and prove non-expansiveness and Markovian convergence of the resulting coupled scheme. Finally, we show that synchronous exact updates are gain-independent at the quotient-law level, isolating a structural contrast between ideal quotient distributions and practical fixed-grid categorical representations.

强化学习分布式学习平均奖励

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。