用时窗逻辑约束强化学习,让机器人在任务中更安全
Hyperproperty-Constrained Secure Reinforcement Learning
- 用时窗逻辑建模安全规则,动态软最大值算法学安全策略
- 在搬运任务中验证,性能优于两种基线方法
- 适合需保障隐私与并发安全的机器人系统
时窗时间逻辑(HyperTWTL)是一种专门用于机器人应用中紧凑表达安全、隐私和并发性质的形式化规范语言。本文聚焦于基于HyperTWTL约束的安全强化学习(SecRL)。尽管时序逻辑约束的安全强化学习(SRL)已有研究,但利用超性质实现安全感知强化学习仍存在显著空白。针对智能体动态作为马尔可夫决策过程(MDP),并将隐私/安全约束形式化为HyperTWTL,本文提出一种基于动态玻尔兹曼软最大值强化学习的方法,以学习满足HyperTWTL约束的安全最优策略。通过一个取送任务的机器人案例研究,验证了该方法的有效性与可扩展性,并与两种基准强化学习算法对比,结果表明本方法表现更优。
原文摘要 · Abstract (English)
Hyperproperties for Time Window Temporal Logic (HyperTWTL) is a domain-specific formal specification language known for its effectiveness in compactly representing security, opacity, and concurrency properties for robotics applications. This paper focuses on HyperTWTL-constrained secure reinforcement learning (SecRL). Although temporal logic-constrained safe reinforcement learning (SRL) is an evolving research problem with several existing literature, there is a significant research gap in exploring security-aware reinforcement learning (RL) using hyperproperties. Given the dynamics of an agent as a Markov Decision Process (MDP) and opacity/security constraints formalized as HyperTWTL, we propose an approach for learning security-aware optimal policies using dynamic Boltzmann softmax RL while satisfying the HyperTWTL constraints. The effectiveness and scalability of our proposed approach are demonstrated using a pick-up and delivery robotic mission case study. We also compare our results with two other baseline RL algorithms, showing that our proposed method outperforms them.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。