用大模型自动生成可组合的奖励机,让RL更易设计且能零样本泛化。
ARM-FM: Automated Reward Machines via Foundation Models for Compositional Reinforcement Learning
- 用大模型从自然语言自动构建奖励机(RM)
- RM状态关联语言嵌入,实现跨任务泛化
- 在多个挑战性环境验证,支持零样本迁移
强化学习(RL)算法对奖励函数设计极为敏感,这仍是限制其广泛应用的核心挑战。本文提出ARM-FM:基于基础模型的自动化奖励机框架,用于强化学习中的可组合奖励设计。该框架利用基础模型(FMs)的高层推理能力,以奖励机(RMs)这一基于自动机的形式化方法来指定RL目标,并通过基础模型自动构建奖励机。奖励机的结构化形式可实现有效的任务分解,而基础模型则使目标规范可通过自然语言表达。具体而言,我们(i)利用基础模型从自然语言规范自动生成奖励机;(ii)将语言嵌入关联至每个奖励机状态,以实现跨任务泛化;(iii)在一系列具有挑战性的环境中提供了实证证据,证明了ARM-FM的有效性,包括零样本泛化能力。
原文摘要 · Abstract (English)
Reinforcement learning (RL) algorithms are highly sensitive to reward function specification, which remains a central challenge limiting their broad applicability. We present ARM-FM: Automated Reward Machines via Foundation Models, a framework for automated, compositional reward design in RL that leverages the high-level reasoning capabilities of foundation models (FMs). Reward machines (RMs) -- an automata-based formalism for reward specification -- are used as the mechanism for RL objective specification, and are automatically constructed via the use of FMs. The structured formalism of RMs yields effective task decompositions, while the use of FMs enables objective specifications in natural language. Concretely, we (i) use FMs to automatically generate RMs from natural language specifications; (ii) associate language embeddings with each RM automata-state to enable generalization across tasks; and (iii) provide empirical evidence of ARM-FM's effectiveness in a diverse suite of challenging environments, including evidence of zero-shot generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。