用大模型自动设计机器人行走策略,减少人工干预。
AnyBipe: An End-to-End Framework for Training and Deploying Bipedal Robots Guided by Large Language Models
- 大模型生成奖励函数,自动优化控制策略。
- 在仿真中训练后直接部署到真实机器人,无需手动调参。
- 适合希望减少人工调试的机器人研发团队。
训练和部署机器人强化学习策略,尤其在完成特定任务时面临巨大挑战。近期研究探索了多种奖励函数设计、训练技术、仿真到现实的迁移方法及性能分析,但仍需大量人工介入。本文提出一个端到端框架,通过大语言模型(LLMs)引导双足机器人强化学习策略的训练与部署,并在双足机器人上评估其有效性。该框架包含三个相互关联的模块:基于大模型的奖励函数设计模块、利用已有工作的强化学习训练模块,以及仿真到现实的同构评估模块。此设计显著降低对人工输入的需求,仅需必要的仿真与部署平台,可集成人工策略与历史数据。我们详述了各模块构建方式及其相较于传统方法的优势,并展示了该框架能自主开发与优化双足机器人行走控制策略,具备独立运行潜力。
原文摘要 · Abstract (English)
Training and deploying reinforcement learning (RL) policies for robots, especially in accomplishing specific tasks, presents substantial challenges. Recent advancements have explored diverse reward function designs, training techniques, simulation-to-reality (sim-to-real) transfers, and performance analysis methodologies, yet these still require significant human intervention. This paper introduces an end-to-end framework for training and deploying RL policies, guided by Large Language Models (LLMs), and evaluates its effectiveness on bipedal robots. The framework consists of three interconnected modules: an LLM-guided reward function design module, an RL training module leveraging prior work, and a sim-to-real homomorphic evaluation module. This design significantly reduces the need for human input by utilizing only essential simulation and deployment platforms, with the option to incorporate human-engineered strategies and historical data. We detail the construction of these modules, their advantages over traditional approaches, and demonstrate the framework's capability to autonomously develop and refine controlling strategies for bipedal robot locomotion, showcasing its potential to operate independently of human intervention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。