arXiv:2501.00364cs.AIcs.FL2025-01被引 4

用一阶逻辑提升奖励机器表达力,让强化学习更高效可迁移

FORM: Learning Expressive and Transferable First-Order Logic Reward Machines

  • 用一阶逻辑替代命题逻辑标注边,大幅增强表达能力
  • 在复杂任务上成功训练传统方法失败的奖励机器
  • 多智能体协作学习框架加速训练并提升跨任务迁移性

奖励机器(RMs)通过有限状态机有效处理强化学习中的非马尔可夫奖励问题。传统RMs使用命题逻辑公式标注边,受限于其表达力,导致复杂任务需大量状态和边,影响可学习性和可迁移性。为此,我们提出一阶奖励机器(FORM),采用一阶逻辑标注边,实现更紧凑、可迁移的模型。我们设计了新的学习方法与多智能体利用框架,使多个智能体协同学习共享的FORM。实验表明,FORM在可扩展性上显著优于传统RMs:在传统方法失效的任务中仍能有效学习,并因一阶逻辑抽象和多智能体框架,实现学习速度提升和任务迁移性能增强。

原文摘要 · Abstract (English)

Reward machines (RMs) are an effective approach for addressing non-Markovian rewards in reinforcement learning (RL) through finite-state machines. Traditional RMs, which label edges with propositional logic formulae, inherit the limited expressivity of propositional logic. This limitation hinders the learnability and transferability of RMs since complex tasks will require numerous states and edges. To overcome these challenges, we propose First-Order Reward Machines ($\texttt{FORM}$s), which use first-order logic to label edges, resulting in more compact and transferable RMs. We introduce a novel method for $\textbf{learning}$ $\texttt{FORM}$s and a multi-agent formulation for $\textbf{exploiting}$ them and facilitate their transferability, where multiple agents collaboratively learn policies for a shared $\texttt{FORM}$. Our experimental results demonstrate the scalability of $\texttt{FORM}$s with respect to traditional RMs. Specifically, we show that $\texttt{FORM}$s can be effectively learnt for tasks where traditional RM learning approaches fail. We also show significant improvements in learning speed and task transferability thanks to the multi-agent learning framework and the abstraction provided by the first-order language.

强化学习奖励机器一阶逻辑多智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。