用大模型自动生成机器人策略,让机器人会自我优化。
SAS-Prompt: Large Language Models as Numerical Optimizers for Robot Self-Improvement
- 用SAS提示框架让大模型分析历史行为并生成新策略
- 在仿真和真实乒乓球任务中实现策略迭代优化
- 无需额外训练,全靠提示工程完成可解释的策略搜索
我们展示了大型语言模型(LLMs)在机器人策略迭代自改进中的能力。关键发现是:大模型具备内在的(随机)数值优化能力,这一特性可用于可解释的机器人策略搜索。基于此,我们提出SAS Prompt(总结、分析、综合)——一种单一提示,通过结合大模型对历史机器人轨迹的检索、推理与优化能力,实现机器人行为的迭代学习与适应,生成前所未见的新行为。该方法可视为一类完全在大模型内部实现的新型可解释策略搜索方法的早期范例。我们在仿真环境和真实机器人乒乓球任务中评估了该方法的有效性。
原文摘要 · Abstract (English)
We demonstrate the ability of large language models (LLMs) to perform iterative self-improvement of robot policies. An important insight of this paper is that LLMs have a built-in ability to perform (stochastic) numerical optimization and that this property can be leveraged for explainable robot policy search. Based on this insight, we introduce the SAS Prompt (Summarize, Analyze, Synthesize) -- a single prompt that enables iterative learning and adaptation of robot behavior by combining the LLM's ability to retrieve, reason and optimize over previous robot traces in order to synthesize new, unseen behavior. Our approach can be regarded as an early example of a new family of explainable policy search methods that are entirely implemented within an LLM. We evaluate our approach both in simulation and on a real-robot table tennis task. Project website: sites.google.com/asu.edu/sas-llm/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。