用大模型复盘过往经验,让机器人听懂复杂指令自动摆物
Learn from the Past: Language-conditioned Object Rearrangement with Large Language Models
- 基于大模型记忆过往成功案例,推理当前摆放策略
- 零样本适配多种日常物品和自由语言指令
- 可处理长序列操作指令,提升机器人泛化能力
物体重新排列至特定目标状态是协作机器人的重要任务。准确确定物体放置位置是关键挑战,因错位会增加任务复杂度与碰撞风险,影响排列效率。现有方法严重依赖预收集数据集训练模型预测目标位置,因而仅适用于特定指令,限制了其广泛适用性和泛化能力。本文提出一种基于大语言模型(LLM)的灵活语言条件物体重排框架。该方法模拟人类推理,利用过往成功经验作为参考,推断实现当前目标位置的最佳策略。得益于LLM强大的自然语言理解与推理能力,本方法能零样本泛化至多种日常物体及自由形式语言指令。实验表明,该方法可有效执行机器人重排任务,包括涉及长序列指令的情况。
原文摘要 · Abstract (English)
Object manipulation for rearrangement into a specific goal state is a significant task for collaborative robots. Accurately determining object placement is a key challenge, as misalignment can increase task complexity and the risk of collisions, affecting the efficiency of the rearrangement process. Most current methods heavily rely on pre-collected datasets to train the model for predicting the goal position. As a result, these methods are restricted to specific instructions, which limits their broader applicability and generalisation. In this paper, we propose a framework of flexible language-conditioned object rearrangement based on the Large Language Model (LLM). Our approach mimics human reasoning by making use of successful past experiences as a reference to infer the best strategies to achieve a current desired goal position. Based on LLM's strong natural language comprehension and inference ability, our method generalises to handle various everyday objects and free-form language instructions in a zero-shot manner. Experimental results demonstrate that our methods can effectively execute the robotic rearrangement tasks, even those involving long sequences of orders.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。