arXiv:2409.04744cs.LGcs.AI2024-09被引 20

用大模型指导强化学习,让机器人更高效地试错。

Reward Guidance for Reinforcement Learning Tasks Based on Large Language Models: The LMGT Framework

  • 用大模型分析教程文本,动态调整奖励信号以引导探索
  • 在机器人任务中减少50%以上训练样本需求,性能优于基线方法
  • 适合缺乏奖励信号的复杂环境,如真实世界机器人控制

强化学习中的环境转移模型存在固有不确定性,需在探索与利用之间取得精妙平衡,以高效估算预期奖励。在奖励稀疏的场景(如机器人控制)中,这一平衡尤为困难。然而,许多环境具备丰富先验知识,从零开始学习可能冗余。为此,我们提出语言模型引导的奖励调优(LMGT)框架,利用大语言模型(LLMs)所蕴含的广泛先验知识及其处理非标准数据(如维基教程)的能力,通过大模型引导的奖励偏移,有效平衡探索与利用,优化智能体的探索行为并提升样本效率。我们在多个强化学习任务及具身机器人环境Housekeep上进行了严格评估,结果表明LMGT持续优于基线方法,并显著降低训练阶段的计算资源消耗。

原文摘要 · Abstract (English)

The inherent uncertainty in the environmental transition model of Reinforcement Learning (RL) necessitates a delicate balance between exploration and exploitation. This balance is crucial for optimizing computational resources to accurately estimate expected rewards for the agent. In scenarios with sparse rewards, such as robotic control systems, achieving this balance is particularly challenging. However, given that many environments possess extensive prior knowledge, learning from the ground up in such contexts may be redundant. To address this issue, we propose Language Model Guided reward Tuning (LMGT), a novel, sample-efficient framework. LMGT leverages the comprehensive prior knowledge embedded in Large Language Models (LLMs) and their proficiency in processing non-standard data forms, such as wiki tutorials. By utilizing LLM-guided reward shifts, LMGT adeptly balances exploration and exploitation, thereby guiding the agent's exploratory behavior and enhancing sample efficiency. We have rigorously evaluated LMGT across various RL tasks and evaluated it in the embodied robotic environment Housekeep. Our results demonstrate that LMGT consistently outperforms baseline methods. Furthermore, the findings suggest that our framework can substantially reduce the computational resources required during the RL training phase.

强化学习大模型机器人控制奖励设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。