arXiv:2411.01775cs.ROcs.AI2024-11CoRL被引 13

用大模型自动生成机器人训练的渐进式环境课程

Eurekaverse: Environment Curriculum Generation via Large Language Models

  • 用大模型生成代码形式的渐进挑战环境
  • 在四足奔跑训练中实现复杂技能的逐步掌握
  • 自动生成课程可迁移到真实机器人,优于人工设计

近期研究表明,通过在由简单到复杂的环境课程中训练机器人,是学习多种复杂技能的有效策略。然而,设计有效的环境分布课程需要大量专业知识,且每新领域都需重复此过程。我们的核心洞察是:环境常以代码形式自然表达。因此,我们探究是否可通过大语言模型(LLM)生成代码来自动化并实现高效环境课程设计。本文提出Eurekaverse,一种无监督的环境设计算法,利用LLM生成逐步更复杂、多样且可学习的环境,用于技能训练。我们在四足机器人越障训练领域验证了其有效性,自动构建的课程使机器人在仿真中逐步掌握复杂越障技能,并成功迁移至真实世界,表现优于人类设计的手动训练课程。

原文摘要 · Abstract (English)

Recent work has demonstrated that a promising strategy for teaching robots a wide range of complex skills is by training them on a curriculum of progressively more challenging environments. However, developing an effective curriculum of environment distributions currently requires significant expertise, which must be repeated for every new domain. Our key insight is that environments are often naturally represented as code. Thus, we probe whether effective environment curriculum design can be achieved and automated via code generation by large language models (LLM). In this paper, we introduce Eurekaverse, an unsupervised environment design algorithm that uses LLMs to sample progressively more challenging, diverse, and learnable environments for skill training. We validate Eurekaverse's effectiveness in the domain of quadrupedal parkour learning, in which a quadruped robot must traverse through a variety of obstacle courses. The automatic curriculum designed by Eurekaverse enables gradual learning of complex parkour skills in simulation and can successfully transfer to the real-world, outperforming manual training courses designed by humans.

机器人课程设计大模型仿真训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。