arXiv:2608.07433math.OCcs.LG2026-08

用最优传输方法求解带熵正则的线性二次控制,可精确化为有限维微分方程。

Wasserstein Policy Gradient for Entropy-Regularized Linear-Quadratic Control

论文配图:Wasserstein Policy Gradient for Entropy-Regularized Linear-Quadratic Control
图 1 · 摘自论文原文
  • 基于动作空间的最优传输构造策略梯度。
  • 收敛速度指数级,且在熵温度趋近零时仍保持稳定。
  • 适合研究带熵正则的强化学习控制问题的学者。

Wasserstein策略梯度(WPG)通过动作空间中的传输更新状态条件动作律。本文研究带熵正则化的折扣线性二次(LQ)控制问题。贝尔曼验证论证表明,无约束问题具有线性高斯最优策略,且折扣占据加权的状态-逐点Wasserstein梯度与此策略类相切。因此,WPG精确简化为反馈增益与动作协方差的有限维常微分方程。我们证明该方程对所有合法初值全局适定,并指数收敛。对于每个固定LQ问题,当熵温度趋于零时,收敛指数有正极限,且不包含形如\exp(-c/τ)的扰动因子,同时保留了对控制问题条件数的正常依赖。

原文摘要 · Abstract (English)

Wasserstein policy gradient (WPG) updates state-conditional action laws by transport in the action space. We study entropy-regularized discounted linear-quadratic (LQ) control. A Bellman verification argument shows that the unrestricted problem has a linear-Gaussian optimal policy, and the discounted-occupancy-weighted statewise Wasserstein gradient is tangent to this policy class. WPG therefore reduces exactly to a finite-dimensional ODE for the feedback gain and action covariance. We prove that this ODE is globally well posed and converges exponentially from every admissible initialization. For each fixed LQ problem, the exponent has a positive limit as the entropy temperature tends to zero and contains no perturbative factor of the form $\exp(-c/τ)$, while retaining the usual dependence on the conditioning of the control problem.

强化学习最优控制熵正则Wasserstein

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。