arXiv:2506.20061cs.LGcs.CL2025-06被引 2

用大模型自动重标注失败轨迹,提升指令跟随强化学习的效率。

Learning Instruction-Following Policies through Open-Ended Instruction Relabeling with Large Language Models

  • 用大模型从智能体轨迹中自动生成开放式指令
  • 在Craftax上实现更高样本效率和指令覆盖率
  • 适合研究指令跟随与强化学习融合的学者

在强化学习中构建有效的指令跟随策略仍具挑战性,主要受限于对大量人工标注指令数据集的依赖以及稀疏奖励下的学习困难。本文提出一种新方法,利用大语言模型(LLM)从已收集的智能体轨迹中回溯生成开放式的指令。核心思想是通过LLM识别智能体隐式完成的有意义子任务,对失败轨迹进行语义重标注,从而丰富训练数据,显著减少对人工标注的依赖。通过这种开放式的指令重标注,我们高效地学习到一个统一的指令跟随策略,可处理多种任务。我们在具有挑战性的Craftax环境中进行实证评估,结果表明,相比现有最优基线,该方法在样本效率、指令覆盖范围和整体策略性能上均有明显提升。结果凸显了利用LLM引导的开放式指令重标注来增强指令跟随强化学习的有效性。

原文摘要 · Abstract (English)

Developing effective instruction-following policies in reinforcement learning remains challenging due to the reliance on extensive human-labeled instruction datasets and the difficulty of learning from sparse rewards. In this paper, we propose a novel approach that leverages the capabilities of large language models (LLMs) to automatically generate open-ended instructions retrospectively from previously collected agent trajectories. Our core idea is to employ LLMs to relabel unsuccessful trajectories by identifying meaningful subtasks the agent has implicitly accomplished, thereby enriching the agent's training data and substantially alleviating reliance on human annotations. Through this open-ended instruction relabeling, we efficiently learn a unified instruction-following policy capable of handling diverse tasks within a single policy. We empirically evaluate our proposed method in the challenging Craftax environment, demonstrating clear improvements in sample efficiency, instruction coverage, and overall policy performance compared to state-of-the-art baselines. Our results highlight the effectiveness of utilizing LLM-guided open-ended instruction relabeling to enhance instruction-following reinforcement learning.

指令跟随强化学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。