arXiv:2607.17897cs.LG2026-07

提出在Cramér几何下分布软贝尔曼算子的收缩性质,为最大熵强化学习提供理论支撑。

Distributional Soft Bellman Operator under the Cramér Geometry

  • 基于累积分布函数的Cramér度量,证明软贝尔曼算子是√γ-收缩的
  • 在统一一阶矩条件下,实现分布估计的唯一不动点和收敛迭代
  • 构建谱域等价表示,助力近似评估器与损失函数设计

分布软策略迭代(DSPI)将分布强化学习与最大熵控制结合,其策略评估步骤由作用于熵正则回报的分布软贝尔曼算子驱动。本文聚焦于基于累积分布函数(CDF)的Cramér几何——一种具有L²结构的概率度量,研究固定策略下的分布软贝尔曼算子是否具备收缩性,从而保证唯一不动点。通过在可接受的CDF域上直接建模,我们定义了CDF层级的分布软贝尔曼算子,证明其为√γ-收缩,并获得对应的唯一不动点及收敛的迭代策略评估方法。该结果表明,有限Cramér域性质源于组合一步奖励熵偏移的统一一阶矩条件,而非对奖励与熵项分别施加有界性假设。进一步通过共轭变换将问题转化为谱域,得到同一决策过程的等价希尔伯特空间表示。这些成果确立了DSPI中策略评估步骤的Cramér几何不动点,为近似评估器、评估误差分析与批评者损失设计提供了理论参考。

原文摘要 · Abstract (English)

Distributional soft policy iteration (DSPI) provides an important framework for combining distributional reinforcement learning (DRL) with maximum-entropy control, in which the policy evaluation step is governed by a distributional soft Bellman operator acting on entropy-regularised returns. Theoretical analysis of such an evaluation step requires a probability metric under which Bellman updates can be controlled, typically by showing that the operator contracts the distance between any two candidate return-distribution estimates. In this paper, we focus on the Cramér geometry, a cumulative distribution function (CDF)-based metric with an $L^2$ structure, and study whether the fixed-policy distributional soft Bellman operator has this contraction property and hence a unique fixed point under this metric. Working directly on an admissible CDF field domain, we formulate the CDF-level distributional soft Bellman operator, prove that it is a $\sqrtγ$-contraction, and obtain the corresponding unique fixed point together with convergent iterative policy evaluation. The CDF formulation also shows that this finite-Cramér-domain property follows from a uniform first-moment condition on the combined one-step reward entropy shift, rather than from separate uniform boundedness assumptions on the reward and entropy terms. We then transport the same evaluation problem to the spectral domain by conjugation, obtaining an equivalent Hilbert-space representation of the same decision process. Taken together, these results identify the Cramér-geometric Bellman fixed point associated with the policy-evaluation step of DSPI, providing a reference point for studying approximate critics, evaluation error, and critic-loss design in DSPI-style algorithms.

强化学习分布学习贝尔曼算子最大熵

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。