arXiv:2503.05563cs.LGmath.OC2025-03中稿 · RLDM 2025被引 2

提出连续时间分布强化学习的可计算近似方法,解决返回分布求解难题。

Tractable Representations for Convergent Approximation of Distributional HJB Equations

  • 通过统计量映射的拓扑性质,建立分布化HJB方程的近似求解条件。
  • 证明分位数表示法满足该性质,可高效逼近连续时间分布强化学习解。
  • 为风险敏感决策的连续时间强化学习提供理论支持,适合研究者参考。

在强化学习中,决策策略的长期行为基于平均回报评估。分布强化学习(Distributional RL)提供了学习回报分布的技术,可提供更多统计信息以评估策略并融入风险敏感考量。当时间无法自然划分为离散增量时,研究者关注连续时间强化学习(CTRL),其中智能体状态与决策持续演化。在此设定下,哈密顿-雅可比-贝尔曼(HJB)方程是期望回报的标准刻画,已有多种求解方法。然而,分布强化学习在连续时间场景的研究仍处于起步阶段。近期工作建立了分布化HJB(DHJB)方程,首次对连续时间强化学习中的回报分布进行了刻画。这些方程及其解难以精确求解与表示,需发展新型近似技术。本文迈出关键一步,确立了在何种参数化返回分布条件下,可近似求解DHJB方程。特别地,我们证明:若分布强化学习算法所学统计量与对应分布间的映射具有特定拓扑性质,则统计量的近似即可导致DHJB解的接近近似。具体而言,我们验证了分布强化学习中常见的分位数表示法满足该拓扑性质,从而为连续时间分布强化学习提供了一个高效的近似算法。

原文摘要 · Abstract (English)

In reinforcement learning (RL), the long-term behavior of decision-making policies is evaluated based on their average returns. Distributional RL has emerged, presenting techniques for learning return distributions, which provide additional statistics for evaluating policies, incorporating risk-sensitive considerations. When the passage of time cannot naturally be divided into discrete time increments, researchers have studied the continuous-time RL (CTRL) problem, where agent states and decisions evolve continuously. In this setting, the Hamilton-Jacobi-Bellman (HJB) equation is well established as the characterization of the expected return, and many solution methods exist. However, the study of distributional RL in the continuous-time setting is in its infancy. Recent work has established a distributional HJB (DHJB) equation, providing the first characterization of return distributions in CTRL. These equations and their solutions are intractable to solve and represent exactly, requiring novel approximation techniques. This work takes strides towards this end, establishing conditions on the method of parameterizing return distributions under which the DHJB equation can be approximately solved. Particularly, we show that under a certain topological property of the mapping between statistics learned by a distributional RL algorithm and corresponding distributions, approximation of these statistics leads to close approximations of the solution of the DHJB equation. Concretely, we demonstrate that the quantile representation common in distributional RL satisfies this topological property, certifying an efficient approximation algorithm for continuous-time distributional RL.

分布强化学习连续时间最优控制数学分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。