arXiv:2412.06877cs.CLcs.AI2024-12ICML被引 4

用大模型增强低质数据,让机器人学会听懂指令并完成新任务

The Synergy of LLMs & RL Unlocks Offline Learning of Generalizable Language-Conditioned Policies with Low-fidelity Data

  • 用大模型自动给无标注数据加语义标签,提升数据价值
  • 在符号环境中实现低样本下稳健的语言指令执行,超越传统RL方法
  • 适合想用少量数据训练通用语言控制策略的研究者

开发能够执行复杂多步决策任务的自主智能体仍是一大挑战,尤其在真实场景中标签数据稀缺且无法实时实验的情况下。现有强化学习方法常难以泛化到未见目标和状态,限制了应用范围。本文提出TEDUO,一种面向符号环境的离线语言条件策略学习新范式。不同于传统方法,TEDUO利用现成的未标注数据集,通过大语言模型(LLMs)双重作用:首先作为自动化工具为离线数据集添加更丰富的标注信息;其次作为可泛化的指令遵循代理。实验证明,TEDUO实现了高效的数据利用,训练出鲁棒的语言条件策略,在任务覆盖范围上超越传统强化学习框架或仅使用大模型的方案。

原文摘要 · Abstract (English)

Developing autonomous agents capable of performing complex, multi-step decision-making tasks specified in natural language remains a significant challenge, particularly in realistic settings where labeled data is scarce and real-time experimentation is impractical. Existing reinforcement learning (RL) approaches often struggle to generalize to unseen goals and states, limiting their applicability. In this paper, we introduce TEDUO, a novel training pipeline for offline language-conditioned policy learning in symbolic environments. Unlike conventional methods, TEDUO operates on readily available, unlabeled datasets and addresses the challenge of generalization to previously unseen goals and states. Our approach harnesses large language models (LLMs) in a dual capacity: first, as automatization tools augmenting offline datasets with richer annotations, and second, as generalizable instruction-following agents. Empirical results demonstrate that TEDUO achieves data-efficient learning of robust language-conditioned policies, accomplishing tasks beyond the reach of conventional RL frameworks or out-of-the-box LLMs alone.

语言控制离线学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。