用信号时序逻辑增强奖励机器,让强化学习更高效完成复杂任务
On Tackling Complex Tasks with Reward Machines and Signal Temporal Logics
- 结合信号时序逻辑与奖励机器,动态生成任务相关奖励信号
- 在微电网、倒立摆和高速路场景中成功实现复杂行为控制
- 支持在线监控,适合需严格满足约束的工业级控制任务
我们提出一种基于强化学习的控制设计框架,用于处理复杂任务。该方法将奖励机器(Reward Machines, RM)扩展为支持信号时序逻辑(Signal Temporal Logic, STL)公式,用于事件生成。STL不仅提升了复杂任务奖励的表达效率,还能引导训练过程收敛到满足指定要求的行为。我们还实现了基于STL在线监测算法的框架变体。通过三个案例研究(微电网、倒立摆、高速公路环境)展示了该方法在非平凡任务上的有效性。
原文摘要 · Abstract (English)
We propose a Reinforcement Learning (RL) based control design framework for handling complex tasks. The approach extends the concept of Reward Machines (RM) with Signal Temporal Logic (STL) formulas that can be used for event generation. The use of STL allows not only a more efficient representation of rewards for complex tasks but also guiding the training process to converge towards behaviors satisfying specified requirements. We also propose an implementation of the framework that leverages the STL online monitoring algorithms. We illustrate the framework with three case studies (minigrid, cart-pole and high-way environments) with non-trivial tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。