让智能体在学习中自动平衡奖励与约束满足,保障任务安全可靠。
Probabilistic Satisfaction of Temporal Logic Constraints in Reinforcement Learning via Adaptive Policy-Switching
- 通过动态切换学习与约束策略,自适应调整行为模式。
- 能有效估计并维持任务约束的满足概率,确保长期合规性。
- 适合对安全性要求高的强化学习应用,如自动驾驶、机器人控制。
约束强化学习(CRL)在传统强化学习框架中引入了额外约束,以满足特定任务需求或限制条件。本文研究一种新型CRL问题:智能体在最大化累积奖励的同时,需在整个学习过程中保证某种时序逻辑约束的满足概率达到预定水平。为此,提出一种新框架,通过在纯学习(奖励最大化)和约束满足策略间动态切换来实现。该框架基于前期试验结果估计约束满足概率,并自适应调整切换概率。理论上证明了算法正确性,并通过全面仿真验证了其性能。
原文摘要 · Abstract (English)
Constrained Reinforcement Learning (CRL) is a subset of machine learning that introduces constraints into the traditional reinforcement learning (RL) framework. Unlike conventional RL which aims solely to maximize cumulative rewards, CRL incorporates additional constraints that represent specific mission requirements or limitations that the agent must comply with during the learning process. In this paper, we address a type of CRL problem where an agent aims to learn the optimal policy to maximize reward while ensuring a desired level of temporal logic constraint satisfaction throughout the learning process. We propose a novel framework that relies on switching between pure learning (reward maximization) and constraint satisfaction. This framework estimates the probability of constraint satisfaction based on earlier trials and properly adjusts the probability of switching between learning and constraint satisfaction policies. We theoretically validate the correctness of the proposed algorithm and demonstrate its performance through comprehensive simulations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。