arXiv:2602.02814math.OCcs.RO2026-02被引 1

为非线性部分可观系统设计通用近似策略并给出性能上限

Sub-optimality bounds for certainty equivalent policies in partially observed systems

  • 用任意状态估计替代最优估计算法构造近似控制策略
  • 在平滑系统中推导出策略次优性的理论上限
  • 适用于需简化计算的复杂控制系统设计

本文推广了随机控制中的确定性等价原理。经典确定性等价原理对线性系统输出反馈与二次代价的情形可解释为:每个时刻的最优动作是将最优状态反馈策略作用于状态的最小均方误差(MMSE)估计值所得。受此启发,本文考虑针对一般(非线性)部分可观随机系统的确定性等价策略,允许使用任意状态估计而非仅限于MMSE估计。在此类设置下,确定性等价策略并非最优。对于代价函数和动态模型在适当意义上光滑的系统,本文推导出确定性等价策略次优性的上界。通过多个例子展示了结果的有效性。

原文摘要 · Abstract (English)

In this paper, we present a generalization of the certainty equivalence principle of stochastic control. One interpretation of the classical certainty equivalence principle for linear systems with output feedback and quadratic costs is as follows: the optimal action at each time is obtained by evaluating the optimal state-feedback policy of the stochastic linear system at the minimum mean square error (MMSE) estimate of the state. Motivated by this interpretation, we consider certainty equivalent policies for general (non-linear) partially observed stochastic systems that allow for any state estimate rather than restricting to MMSE estimates. In such settings, the certainty equivalent policy is not optimal. For models where the cost and the dynamics are smooth in an appropriate sense, we derive upper bounds on the sub-optimality of certainty equivalent policies. We present several examples to illustrate the results.

控制理论随机优化系统辨识

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。