用上下文学习让大模型迭代优化机械臂抓取策略,数据少也能高效适应。
In-Context Iterative Policy Improvement for Dynamic Manipulation
- 基于上下文学习,通过历史交互逐步调整参数化策略。
- 在仿真和真实机器人上均实现低数据场景下优于传统方法的性能。
- 适合对少样本动态操作有需求的研究者或工业应用开发者。
基于互联网规模语言数据训练的注意力架构在逻辑推理和文本理解等语言任务中展现出顶尖的推理能力。这些大语言模型(LLMs)还具备通过上下文学习实现少样本预测的能力,即利用提示中的输入输出示例泛化到新输入。该能力不仅限于标准语言任务,还可推广至一般模式的少样本学习。本文探讨将预训练语言模型的上下文学习应用于动态操作任务。动态操作面临高维、复杂动力学和部分可观测性等关键挑战。为此,我们采用迭代方法,将上下文学习问题建模为根据先前交互预测参数化策略的调整。我们在多个仿真任务及物理机器人实验中验证,使用上下文学习在低数据条件下表现优于其他方法。本工作的视频摘要及实验演示可访问:https://youtu.be/2inxpdrq74U?si=dAdDYsUEr25nZvRn。
原文摘要 · Abstract (English)
Attention-based architectures trained on internet-scale language data have demonstrated state of the art reasoning ability for various language-based tasks, such as logic problems and textual reasoning. Additionally, these Large Language Models (LLMs) have exhibited the ability to perform few-shot prediction via in-context learning, in which input-output examples provided in the prompt are generalized to new inputs. This ability furthermore extends beyond standard language tasks, enabling few-shot learning for general patterns. In this work, we consider the application of in-context learning with pre-trained language models for dynamic manipulation. Dynamic manipulation introduces several crucial challenges, including increased dimensionality, complex dynamics, and partial observability. To address this, we take an iterative approach, and formulate our in-context learning problem to predict adjustments to a parametric policy based on previous interactions. We show across several tasks in simulation and on a physical robot that utilizing in-context learning outperforms alternative methods in the low data regime. Video summary of this work and experiments can be found https://youtu.be/2inxpdrq74U?si=dAdDYsUEr25nZvRn.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。