arXiv:2511.08419eess.SYcs.LG2025-11

用平均奖励MDP方法,为随机控制系统提供可计算的安全保障。

Probabilistic Safety Guarantee for Stochastic Control Systems Using Average Reward MDPs

  • 将安全约束转化为平均奖励MDP问题求解
  • 在双积分器和倒立摆系统上收敛更快、效果更优
  • 适合需要高置信度安全控制的工程应用

随机控制系统的安全性旨在计算满足预设操作约束的策略,以高置信度应对状态变量的不确定演化。由于状态变量的不可预测性,传统控制方法难以满足约束要求。为此,本文提出一种新算法,通过将安全目标转化为标准平均奖励马尔可夫决策过程(Average Reward MDP)目标,实现对有限状态集上安全水平的计算。该转化使我们能够使用线性规划等标准技术来计算和分析安全策略。我们在双积分器(Double Integrator)和倒立摆(Inverted Pendulum)系统上进行了数值验证,结果表明,平均奖励MDP解法比最小折扣奖励解法更具全面性,收敛速度更快,且策略质量更高。

原文摘要 · Abstract (English)

Safety in stochastic control systems, which are subject to random noise with a known probability distribution, aims to compute policies that satisfy predefined operational constraints with high confidence throughout the uncertain evolution of the state variables. The unpredictable evolution of state variables poses a significant challenge for meeting predefined constraints using various control methods. To address this, we present a new algorithm that computes safe policies to determine the safety level across a finite state set. This algorithm reduces the safety objective to the standard average reward Markov Decision Process (MDP) objective. This reduction enables us to use standard techniques, such as linear programs, to compute and analyze safe policies. We validate the proposed method numerically on the Double Integrator and the Inverted Pendulum systems. Results indicate that the average-reward MDPs solution is more comprehensive, converges faster, and offers higher quality compared to the minimum discounted-reward solution.

随机控制MDP安全控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。