arXiv:2606.17762math.OCcs.AI2026-06

研究离散时间庞特里亚金系统中终端奖励扰动的衰减规律,发现其影响随时间呈指数级减弱。

Horizon-Uniform Sensitivity and Decay of Terminal Reward Perturbations in Discrete-Time Pontryagin Systems

  • 基于加权范数的压缩映射证明解的存在唯一性与均匀稳定性。
  • 终端奖励扰动对初始控制和状态梯度的影响以 $O(e^{-α_{\rm ter} T})$ 衰减。
  • 适用于线性二次型系统,可验证最优反馈增益以 $O(e^{-2γT})$ 收敛。

我们研究有限时域离散时间庞特里亚金系统在稳态极值附近的局部平稳解。假设控制的站位方程正则,约化状态-协态映射双曲,端点条件满足关于稳定与不稳定子空间的缩放横截条件,则线性化边值问题的逆存在,格林估计在时域上一致。格林核分离内部衰减与端点条件引起的两次反射。当 $x_0=x_{\rm in}$ 且 $p_T=r_x(x_T,y)$ 时,在加权范数下通过压缩映射证明解在与 $T$ 无关的邻域内存在唯一,具有一致Lipschitz估计和逐点二次余项。我们还推导出可接受数据半径及近似轨迹附近的后验存在性与局部唯一性判据。对于这类图边界条件,单侧格林估计表明:终端奖励的扰动会使初始控制和目标函数关于初态的梯度变化为 $O(e^{-α_{\rm ter} T})$,其中 $α_{\rm ter}$ 小于分歧率。在线性二次系统中,若 $A$ 可逆,$(A,B)$ 可镇定,$Q\succ0$,$R\succ0$,且终端海森矩阵非正,对称辛图条件可验证假设,有限时域里卡提矩阵与初始反馈增益以 $O(e^{-2γT})$ 的速率收敛。数值实验验证了这些结论与预测的衰减速率。

原文摘要 · Abstract (English)

We study local stationary solutions of finite-horizon discrete-time Pontryagin systems near a steady extremal. Suppose that the stationarity equation for the control is regular, the reduced state--costate map is hyperbolic, and the endpoint conditions satisfy a scaled transversality condition with respect to the stable and unstable subspaces. Then the linearized boundary-value problem admits an inverse whose Green estimate is uniform in the horizon. The Green kernel separates interior decay from the two reflections induced by the endpoint conditions. For $x_0=x_{\rm in}$ and $p_T=r_x(x_T,y)$, a contraction argument in a weighted norm proves existence and uniqueness in a neighborhood independent of $T$, together with uniform Lipschitz estimates and a pointwise quadratic remainder. We also derive an explicit admissible data radius and an a posteriori criterion for existence and local uniqueness near an approximate trajectory. For these graph boundary conditions, a one-sided Green estimate shows that a perturbation of the terminal reward changes the initial control and the gradient with respect to the initial state of the stationary objective by $O(e^{-α_{\rm ter} T})$ for every $α_{\rm ter}$ below the dichotomy rate. For linear-quadratic systems with invertible $A$, stabilizable $(A,B)$, $Q\succ0$, $R\succ0$, and a nonpositive terminal Hessian, a symplectic graph condition verifies the assumptions, and the finite-horizon Riccati matrix and initial feedback gain converge at rate $O(e^{-2γT})$. Numerical experiments verify the certificates and the predicted decay rates.

最优控制稳定性分析指数衰减离散系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。