提出新方法,在随机环境中同时保证目标可达且避开危险,还能最小化成本。
Stochastic Minimum-Cost Reach-Avoid Reinforcement Learning

- 用概率证书识别可满足约束的状态空间
- 基于收缩的贝尔曼方程实现概率约束下的成本优化
- 适合需要高可靠性与成本控制的机器人决策场景
我们研究随机最小代价可达-避让强化学习,要求智能体在随机环境中以不低于概率 $p$ 满足可达-避让规范,同时最小化累积期望成本。现有安全与约束强化学习方法通常无法在随机环境下联合执行概率性可达-避让约束并优化成本。为此,我们引入可达-避让概率证书(RAPCs),用于识别从哪些状态出发,可达-避让约束是可满足的。基于 RAPCs,我们构建了基于收缩的贝尔曼公式,作为将可达-避让考量融入强化学习的合理替代目标,从而实现概率约束下的成本优化。我们证明了所提算法几乎必然收敛至局部最优策略。在 MuJoCo 模拟器上的实验表明,该方法在成本性能上更优,且可达-避让满足率始终更高。
原文摘要 · Abstract (English)
We study stochastic minimum-cost reach-avoid reinforcement learning, where an agent must satisfy a reach-avoid specification with probability at least $p$ while minimizing expected cumulative costs in stochastic environments. Existing safe and constrained reinforcement learning methods typically fail to jointly enforce probabilistic reach-avoid constraints and optimize cost in the learning setting in stochastic environments. To address this challenge, we introduce reach-avoid probability certificates (RAPCs), which identify states from which stochastic reach-avoid constraints are satisfiable. Building on RAPCs, we develop a contraction-based Bellman formulation that serves as a principled surrogate for integrating reach-avoid considerations into reinforcement learning, enabling cost optimization under probabilistic constraints. We establish almost sure convergence of the proposed algorithms to locally optimal policies with respect to the resulting objective. Experiments in the MuJoCo simulator demonstrate improved cost performance and consistently higher reach-avoid satisfaction rates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。