arXiv:2505.23244cs.LG2025-05被引 2

证明了连续控制中随机与确定性策略梯度等价,统一了两类方法的理论基础。

Equivalence of stochastic and deterministic policy gradients

  • 在含高斯噪声和二次代价的马尔可夫决策过程下,两类梯度完全一致。
  • 构造新MDP使确定性策略等价于原随机策略,控制量为原策略的充分统计量。
  • 建议用状态价值函数近似统一优化,避免复杂的状态-控制值函数建模。

连续控制中的策略梯度已分别针对随机和确定性策略推导。本文研究两者关系:在广泛使用的涉及高斯控制噪声和二次控制代价的马尔可夫决策过程(MDP)家族中,随机与确定性策略梯度、自然梯度及状态价值函数完全相同;而状态-控制价值函数不同。随后,我们提出一种通用方法,可构建一个等价于原随机策略MDP的确定性策略MDP,其控制量为原策略的充分统计量。结果表明,策略梯度方法可通过近似状态价值函数而非状态-控制价值函数实现统一。

原文摘要 · Abstract (English)

Policy gradients in continuous control have been derived for both stochastic and deterministic policies. Here we study the relationship between the two. In a widely-used family of MDPs involving Gaussian control noise and quadratic control costs, we show that the stochastic and deterministic policy gradients, natural gradients, and state value functions are identical; while the state-control value functions are different. We then develop a general procedure for constructing an MDP with deterministic policy that is equivalent to a given MDP with stochastic policy. The controls of this new MDP are the sufficient statistics of the stochastic policy in the original MDP. Our results suggest that policy gradient methods can be unified by approximating state value functions rather than state-control value functions.

强化学习策略梯度连续控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。