arXiv:2411.15891cs.LG2024-11

用游戏规则生成内部动机,提升智能体探索效率。

From Laws to Motivation: Guiding Exploration through Law-Based Reasoning and Rewards

  • 从交互记录中提炼游戏规律,以语言形式表达为内在动机。
  • 在Crafter环境中,强化学习与大模型代理性能均显著提升。
  • 适合研究自主智能体、奖励设计与环境理解的学者。

大型语言模型(LLMs)和强化学习(RL)是构建自主智能体的两种强大方法。然而,由于对游戏环境理解有限,智能体常依赖低效的探索与试错,难以形成长期策略或做出合理决策。本文提出一种方法,通过分析交互记录提取经验,建模游戏环境的底层规律,并将这些经验以语言形式作为内部动机,指导智能体行为。这些经验既可直接用于推理,也可转化为奖励信号用于训练。在Crafter环境中的评估表明,无论是强化学习还是大语言模型代理,均能从这些经验中获益,整体性能得到提升。

原文摘要 · Abstract (English)

Large Language Models (LLMs) and Reinforcement Learning (RL) are two powerful approaches for building autonomous agents. However, due to limited understanding of the game environment, agents often resort to inefficient exploration and trial-and-error, struggling to develop long-term strategies or make decisions. We propose a method that extracts experience from interaction records to model the underlying laws of the game environment, using these experience as internal motivation to guide agents. These experience, expressed in language, are highly flexible and can either assist agents in reasoning directly or be transformed into rewards for guiding training. Our evaluation results in Crafter demonstrate that both RL and LLM agents benefit from these experience, leading to improved overall performance.

智能体强化学习大模型动机引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。