arXiv:2501.12539cs.LGcs.CL2025-01被引 2

用强化学习让语言模型更高效地理解复合指令,少试几步就学会新任务。

Compositional Instruction Following with Language Models and Reinforcement Learning

  • 通过组合式策略和语义解析器,让模型从语言指令中拆解任务
  • 在162个测试任务上,样本量减少且成功率提升至92%(接近理论上限)
  • 适合研究多任务语言理解、高效强化学习的开发者或研究人员

将强化学习与语言理解结合面临挑战:代理需在探索环境中同时学习多种语言控制任务。为此,我们提出一种新方法——组合式强化学习语言智能体(CERLLA)。该方法通过组合式策略表示和基于强化学习与上下文学习训练的语义解析器,降低语言指定任务的样本复杂度。我们在需要函数逼近的环境中评估该方法,验证了其对新任务的组合泛化能力。在162个用于测试组合泛化能力的任务中,该方法显著优于此前最佳的非组合基线,在样本复杂度上表现优异。模型取得更高成功率,学习步数更少;在相同环境步数下,成功率达92%,接近一个理想策略的上限性能,而基线仅达80%。

原文摘要 · Abstract (English)

Combining reinforcement learning with language grounding is challenging as the agent needs to explore the environment while simultaneously learning multiple language-conditioned tasks. To address this, we introduce a novel method: the compositionally-enabled reinforcement learning language agent (CERLLA). Our method reduces the sample complexity of tasks specified with language by leveraging compositional policy representations and a semantic parser trained using reinforcement learning and in-context learning. We evaluate our approach in an environment requiring function approximation and demonstrate compositional generalization to novel tasks. Our method significantly outperforms the previous best non-compositional baseline in terms of sample complexity on 162 tasks designed to test compositional generalization. Our model attains a higher success rate and learns in fewer steps than the non-compositional baseline. It reaches a success rate equal to an oracle policy's upper-bound performance of 92%. With the same number of environment steps, the baseline only reaches a success rate of 80%.

语言理解强化学习组合泛化智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。