arXiv:2506.08121math.OCcs.LG2025-06被引 1

用连续迭代方法同步优化控制策略与价值函数,适用于无限时域随机控制问题。

Continuous Policy and Value Iteration for Stochastic Control Problems and Its Convergence

  • 通过朗之万动力学同时更新策略与价值函数
  • 在哈密顿单调条件下收敛至最优控制
  • 支持分布采样与非凸学习,适合复杂控制场景

我们提出一种连续策略-价值迭代算法,通过朗之万型随机微分方程,在无限时域下同步更新随机控制问题的价值函数近似与最优控制。该框架适用于熵正则化松弛控制问题与经典控制问题。在哈密顿函数单调性条件下,证明了策略改进并实现了收敛至最优控制。利用朗之万型动态沿策略迭代方向进行连续更新,使机器学习中的分布采样与非凸学习技术可同时用于优化价值函数和识别最优控制。

原文摘要 · Abstract (English)

We introduce a continuous policy-value iteration algorithm where the approximations of the value function of a stochastic control problem and the optimal control are simultaneously updated through Langevin-type dynamics. This framework applies to both the entropy-regularized relaxed control problems and the classical control problems, with infinite horizon. We establish policy improvement and demonstrate convergence to the optimal control under the monotonicity condition of the Hamiltonian. By utilizing Langevin-type stochastic differential equations for continuous updates along the policy iteration direction, our approach enables the use of distribution sampling and non-convex learning techniques in machine learning to optimize the value function and identify the optimal control simultaneously.

强化学习随机控制连续优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。