用对偶上升法解决带概率约束的无限时域控制问题,兼顾性能与计算效率。
Learning-Based Stochastic Optimal Control with Infinite-Horizon Probabilistic Constraints

- 将原问题转化为增广状态下的约束MDP,具可加结构
- 证明强对偶性,转为等价无约束问题并收敛至最优可行策略
- 离线学习值函数,显著降低在线控制计算开销
本文研究具有无限时域联合机会约束的随机最优控制问题。通过适当的状态增广,将原问题重构为具有可加结构的成本与约束函数的约束马尔可夫决策过程(CMDP)。我们证明该形式具备强对偶性,从而可在拉格朗日对偶框架下将其转化为等价的无约束问题。为此提出对偶上升算法求解,并证明其收敛至定义在增广状态空间上的确定性马尔可夫策略,该策略既最优又满足约束。为处理连续状态-输入空间,设计专用学习算法在离线训练中逼近值函数,显著降低在线控制阶段的计算复杂度。通过数值实验验证,所提方法在性能和计算效率上均优于在线预测控制方法。
原文摘要 · Abstract (English)
In this paper, we consider stochastic optimal control problems with infinite-horizon joint chance constraints. By means of an appropriate state augmentation, we reformulate the original problem as a constrained Markov decision process, in which both the cost and the constraint function exhibit an additive structure. We then prove that this formulation enjoys strong duality, thereby enabling us to reformulate the problem as an equivalent unconstrained one in the Lagrange dual framework. We propose a dual-ascent algorithm to solve the resulting problem and show that it converges to a deterministic Markov policy defined over the augmented state space that is both optimal and feasible. To accommodate continuous state-input spaces, we propose a dedicated learning algorithm to approximate the value function in an offline training setting, thereby significantly reducing the computational complexity of the online control phase. We then test our approach on a numerical example and demonstrate its effectiveness compared to online predictive control methods in terms of performance and computational complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。