arXiv:2409.03402cs.AIcs.RO2024-09被引 1

用视觉语言模型自动设计强化学习任务,让机器人自主学会复杂技能。

Game On: Towards Language Models as RL Experimenters

  • 用视觉语言模型替代人工完成实验设计、任务分解和技能检索
  • 在机器人控制领域实现数据收集与策略迭代的自动化,提升学习效率
  • 适合对自主强化学习系统感兴趣的开发者与研究者

我们提出一种代理架构,自动化强化学习实验流程中的部分环节,以实现对具身智能体控制领域的自动化掌握。该架构利用视觉语言模型(VLM)承担人类实验者通常需具备的能力,包括实验进展的监控与分析、基于过往成功与失败提出新任务、将任务分解为子任务(技能序列),以及检索并执行相应技能,从而构建自动化课程体系。这是首个提议在完整强化学习实验周期中全程使用VLM的系统。我们提供首个原型,并检验当前模型与技术实现所需自动化水平的可行性。采用未微调的标准Gemini模型,为语言条件化的演员-评论家算法生成技能课程,引导数据收集以辅助学习新技能。实验表明,此类数据对学习和迭代优化控制策略有效。进一步验证了系统构建持续增长的技能库及评估训练进度的能力,结果令人鼓舞,表明该架构可能成为具身智能体全自动掌握任务与领域的一种可行方案。

原文摘要 · Abstract (English)

We propose an agent architecture that automates parts of the common reinforcement learning experiment workflow, to enable automated mastery of control domains for embodied agents. To do so, it leverages a VLM to perform some of the capabilities normally required of a human experimenter, including the monitoring and analysis of experiment progress, the proposition of new tasks based on past successes and failures of the agent, decomposing tasks into a sequence of subtasks (skills), and retrieval of the skill to execute - enabling our system to build automated curricula for learning. We believe this is one of the first proposals for a system that leverages a VLM throughout the full experiment cycle of reinforcement learning. We provide a first prototype of this system, and examine the feasibility of current models and techniques for the desired level of automation. For this, we use a standard Gemini model, without additional fine-tuning, to provide a curriculum of skills to a language-conditioned Actor-Critic algorithm, in order to steer data collection so as to aid learning new skills. Data collected in this way is shown to be useful for learning and iteratively improving control policies in a robotics domain. Additional examination of the ability of the system to build a growing library of skills, and to judge the progress of the training of those skills, also shows promising results, suggesting that the proposed architecture provides a potential recipe for fully automated mastery of tasks and domains for embodied agents.

强化学习具身智能自动化实验视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。