用时序逻辑约束强化学习,让机器人在动态环境中安全完成复杂任务
Shielded Reinforcement Learning Under Dynamic Temporal Logic Constraints
- 结合控制屏障函数与无模型强化学习,实时保障时序逻辑任务
- 支持动态目标访问等复杂时空约束,不限于传统安全规则
- 适合需要高可靠性的自主系统,如无人机、自动驾驶
强化学习在机器人应用中展现潜力,但实际部署受限于安全与操作约束。近年来,安全强化学习聚焦于学习过程中施加安全限制,然而真实系统常需更复杂的约束,如周期性充电或对特定区域的限时访问。这类时空任务在学习阶段仍难实现。信号时序逻辑(STL)是一种用于描述实值信号时序特性的形式语言,可表达此类复杂任务。本文提出一种框架,结合序列控制屏障函数与无模型强化学习,确保学习过程中满足给定的STL任务。该方法超越传统安全约束,能处理涉及未知轨迹动态目标的丰富STL规范。通过多种仿真验证了框架的有效性。
原文摘要 · Abstract (English)
Reinforcement Learning (RL) has shown promise in various robotics applications, yet its deployment on real systems is still limited due to safety and operational constraints. The safe RL field has gained considerable attention in recent years, which focuses on imposing safety constraints throughout the learning process. However, real systems often require more complex constraints than just safety, such as periodic recharging or time-bounded visits to specific regions. Imposing such spatio-temporal tasks during learning still remains a challenge. Signal Temporal Logic (STL) is a formal language for specifying temporal properties of real-valued signals and provides a way to express such complex tasks. In this paper, we propose a framework that leverages sequential control barrier functions and model-free RL to ensure that the given STL tasks are satisfied throughout the learning process. Our method extends beyond traditional safety constraints by enforcing rich STL specifications, which can involve visits to dynamic targets with unknown trajectories. We also demonstrate the effectiveness of our framework through various simulations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。