arXiv:2410.01212cs.LG2024-10被引 3

提出新算法,让机器人在复杂任务中高概率满足安全约束。

Absolute State-wise Constrained Policy Optimization: High-Probability State-wise Constraints Satisfaction

  • 基于高概率约束设计新策略优化算法,无需强假设
  • 在连续控制任务中显著优于现有方法,约束违反率更低
  • 适合自动驾驶、机器人操控等对安全性要求高的场景

在真实世界应用中,如自动驾驶和机器人操作,确保状态级安全约束至关重要。然而,现有安全强化学习方法要么仅在期望层面约束,无法排除安全违规风险;要么要求强假设下硬性约束,不具实用性。本文洞察到,在无模型设置下难以保证硬性状态约束,但可通过高概率方式实现安全。为此,提出绝对状态约束策略优化(ASCPO),一种通用型策略搜索算法,能在随机系统中高概率满足状态约束。通过在多种机器人运动任务中训练神经网络策略,验证了该方法的有效性。结果表明,ASCPO在挑战性连续控制任务中显著优于现有方法,大幅降低约束违反风险,展现出在实际应用中的巨大潜力。

原文摘要 · Abstract (English)

Enforcing state-wise safety constraints is critical for the application of reinforcement learning (RL) in real-world problems, such as autonomous driving and robot manipulation. However, existing safe RL methods only enforce state-wise constraints in expectation or enforce hard state-wise constraints with strong assumptions. The former does not exclude the probability of safety violations, while the latter is impractical. Our insight is that although it is intractable to guarantee hard state-wise constraints in a model-free setting, we can enforce state-wise safety with high probability while excluding strong assumptions. To accomplish the goal, we propose Absolute State-wise Constrained Policy Optimization (ASCPO), a novel general-purpose policy search algorithm that guarantees high-probability state-wise constraint satisfaction for stochastic systems. We demonstrate the effectiveness of our approach by training neural network policies for extensive robot locomotion tasks, where the agent must adhere to various state-wise safety constraints. Our results show that ASCPO significantly outperforms existing methods in handling state-wise constraints across challenging continuous control tasks, highlighting its potential for real-world applications.

强化学习安全约束连续控制高概率保证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。