用语义标注的自动机提升多任务强化学习的指令理解能力
Semantically Labelled Automata for Multi-Task Reinforcement Learning with LTL Instructions
- 将LTL逻辑公式转为带语义的自动机,状态蕴含丰富结构信息
- 在多个场景中表现超越现有方法,可处理复杂未见任务规范
- 适合需要精准理解时序指令的智能体系统设计者
我们研究多任务强化学习,即让智能体学习一个通用策略以适应任意可能的、甚至未见过的任务。任务以线性时序逻辑(LTL)公式描述,这类形式化语言常用于系统性质建模,近年也成功应用于强化学习。本文提出一种新型任务嵌入技术,利用新一代语义LTL到自动机的转换方法(原用于时序合成)。生成的语义标注自动机在每个状态中包含丰富、结构化的信息,使得我们能够:(i) 实时高效计算自动机;(ii) 提取可用于条件化策略的表达性强的任务嵌入;(iii) 自然支持完整的LTL表达能力。在多种环境中的实验表明,该方法达到当前最优性能,并能扩展到现有方法失效的复杂规范场景。
原文摘要 · Abstract (English)
We study multi-task reinforcement learning (RL), a setting in which an agent learns a single, universal policy capable of generalising to arbitrary, possibly unseen tasks. We consider tasks specified as linear temporal logic (LTL) formulae, which are commonly used in formal methods to specify properties of systems, and have recently been successfully adopted in RL. In this setting, we present a novel task embedding technique leveraging a new generation of semantic LTL-to-automata translations, originally developed for temporal synthesis. The resulting semantically labelled automata contain rich, structured information in each state that allow us to (i) compute the automaton efficiently on-the-fly, (ii) extract expressive task embeddings used to condition the policy, and (iii) naturally support full LTL. Experimental results in a variety of domains demonstrate that our approach achieves state-of-the-art performance and is able to scale to complex specifications where existing methods fail.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。