arXiv:2412.16561math.OCcs.LG2024-12被引 7

无需系统模型,用数据学习随机系统的安全到达控制策略

A learning-based approach to stochastic optimal control under reach-avoid constraint

  • 通过状态扩展将非马尔可夫问题转为马尔可夫决策过程
  • 在有限时域内以高概率满足安全区与目标区约束
  • 基于轨迹数据的梯度方法,适合无模型强化学习场景

我们提出一种无模型方法,用于优化受制于可达-避障约束的随机马尔可夫系统。具体而言,状态轨迹必须在有限时间窗内保持在安全集内并抵达目标集。由于约束具有时变性,我们证明该约束随机控制问题的最优策略通常是非马尔可夫的,从而增加计算复杂度。为应对这一挑战,我们采用arXiv:2402.19360中的状态增强技术,将问题重构为扩展状态空间上的约束马尔可夫决策过程(CMDP),使最优策略可被搜索为马尔可夫策略,避免了非马尔可夫策略带来的复杂性。为在无系统模型且仅使用轨迹数据的情况下学习最优策略,我们开发了一种对数障碍策略梯度方法。在合理假设下,我们证明策略参数能收敛至最优参数,同时确保系统轨迹以高概率满足随机可达-避障约束。

原文摘要 · Abstract (English)

We develop a model-free approach to optimally control stochastic, Markovian systems subject to a reach-avoid constraint. Specifically, the state trajectory must remain within a safe set while reaching a target set within a finite time horizon. Due to the time-dependent nature of these constraints, we show that, in general, the optimal policy for this constrained stochastic control problem is non-Markovian, which increases the computational complexity. To address this challenge, we apply the state-augmentation technique from arXiv:2402.19360, reformulating the problem as a constrained Markov decision process (CMDP) on an extended state space. This transformation allows us to search for a Markovian policy, avoiding the complexity of non-Markovian policies. To learn the optimal policy without a system model, and using only trajectory data, we develop a log-barrier policy gradient approach. We prove that under suitable assumptions, the policy parameters converge to the optimal parameters, while ensuring that the system trajectories satisfy the stochastic reach-avoid constraint with high probability.

强化学习随机控制约束优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。