让强化学习模型学会用新符号理解复杂任务指令
PlatoLTL: Learning to Generalize Across Symbols in LTL Instructions for Multi-Task RL
- 将符号视为可参数化的原子谓词,共享语义结构
- 在未见过的符号组合下实现零样本泛化,成功率超基准30%
- 适合需要灵活应对新任务描述的机器人系统研究者
多任务强化学习的核心挑战是训练出能执行训练中未见任务的通用策略。为促进这种泛化,线性时序逻辑(LTL)作为形式化表达具有时间结构的任务的有力工具被引入。现有方法虽可在不同LTL规范间泛化,但无法处理未见过的命题词汇(即“符号”),这些符号用于描述LTL中的高层事件。我们提出PlatoLTL,一种新方法,使策略不仅能组合式地泛化于LTL结构,还能参数化地泛化于命题。我们将命题建模为原子谓词的参数化实例,使策略能学习相关命题间的共享结构。我们设计了一种新型架构,嵌入并组合参数化命题以表示LTL公式,并在多个挑战性环境中展示了零样本泛化能力。
原文摘要 · Abstract (English)
A central challenge in multi-task reinforcement learning (RL) is to train generalist policies capable of performing tasks not seen during training. To facilitate such generalization, linear temporal logic (LTL) has emerged as a powerful formalism for specifying structured, temporally extended tasks to RL agents. While existing approaches to LTL-guided multi-task RL demonstrate generalization across LTL specifications, they are unable to generalize to unseen vocabularies of propositions (or "symbols"), which describe high-level events in LTL. We present PlatoLTL, a novel approach that enables policies to zero-shot generalize not only compositionally across LTL structures, but also parametrically across propositions. We model propositions as parameterized instances of atomic predicates, allowing policies to learn shared structure across related propositions. We propose a novel architecture that embeds and composes parameterized propositions to represent LTL formulae, and demonstrate zero-shot generalization in a range of challenging environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。