arXiv:2507.15082cs.LGcs.AI2025-07

考虑价值函数梯度不确定性的鲁棒控制新方法,改进了强化学习中的稳定性问题。

Robust Control with Gradient Uncertainty

  • 构建对抗性动态博弈模型,同时扰动系统和价值函数梯度。
  • 在线性二次情形下证明经典二次假设在梯度不确定时失效。
  • 提出GURAC算法并验证其在训练稳定上的有效性,适合强化学习应用。

我们提出一种鲁棒控制理论的新扩展,明确处理价值函数梯度的不确定性,这类不确定性在强化学习等使用函数近似的问题中普遍存在。通过构建一个零和动态博弈,其中对手同时扰动系统动力学和价值函数梯度,导出了一个新的高度非线性偏微分方程:含梯度不确定性的哈密顿-雅可比-贝尔曼-伊斯艾克斯方程(GU-HJBI)。在统一椭圆条件下,我们通过证明粘性解的比较原理,建立了该方程的适定性。对线性二次(LQ)情形的分析揭示关键洞见:任何非零梯度不确定性均使经典二次价值函数假设失效,从根本上改变问题结构。形式摄动分析刻画了价值函数的非多项式修正及最优控制律的非线性特性,并通过数值研究加以验证。最后,我们提出一种新型梯度不确定性鲁棒的演员-评论家算法(GURAC),并通过实证研究展示了其在稳定训练方面的有效性。本工作为鲁棒控制开辟了新方向,对函数近似广泛应用的领域(如强化学习、计算金融)具有重要意义。

原文摘要 · Abstract (English)

We introduce a novel extension to robust control theory that explicitly addresses uncertainty in the value function's gradient, a form of uncertainty endemic to applications like reinforcement learning where value functions are approximated. We formulate a zero-sum dynamic game where an adversary perturbs both system dynamics and the value function gradient, leading to a new, highly nonlinear partial differential equation: the Hamilton-Jacobi-Bellman-Isaacs Equation with Gradient Uncertainty (GU-HJBI). We establish its well-posedness by proving a comparison principle for its viscosity solutions under a uniform ellipticity condition. Our analysis of the linear-quadratic (LQ) case yields a key insight: we prove that the classical quadratic value function assumption fails for any non-zero gradient uncertainty, fundamentally altering the problem structure. A formal perturbation analysis characterizes the non-polynomial correction to the value function and the resulting nonlinearity of the optimal control law, which we validate with numerical studies. Finally, we bridge theory to practice by proposing a novel Gradient-Uncertainty-Robust Actor-Critic (GURAC) algorithm, accompanied by an empirical study demonstrating its effectiveness in stabilizing training. This work provides a new direction for robust control, holding significant implications for fields where function approximation is common, including reinforcement learning and computational finance.

鲁棒控制强化学习函数近似优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。