arXiv:2409.11863cs.ROcs.AI2024-09ICRA被引 3

融合触觉力信息提升大模型对复杂操作任务的理解与规划能力

LEMMo-Plan: LLM-Enhanced Learning from Multi-Modal Demonstration for Planning Sequential Contact-Rich Manipulation Tasks

  • 用触觉和力矩数据增强大模型对多模态示范的理解
  • 在真实机器人上实现两种序列操作任务的高效规划
  • 适合研究人机协作与复杂抓取的开发者参考

大语言模型(LLMs)在长时序操作任务规划中受到关注。为提升生成计划的有效性,视觉演示和在线视频被广泛用于引导规划过程。然而,对于涉及细微动作但富含接触交互的操作任务,仅依赖视觉感知可能不足以让大模型充分理解示范内容。此外,视觉数据对力相关参数和状态的信息有限,而这些对真实机器人执行至关重要。本文提出一种上下文学习框架,将人类示范中的触觉与力矩信息融入,以增强大模型生成新任务场景计划的能力。我们设计了一个自举推理流程,逐步将各模态信息整合进综合任务计划,并以此作为新任务配置下的规划参考。在两个不同序列操作任务的真实世界实验中,验证了该框架能有效提升大模型对多模态示范的理解能力,显著改善整体规划性能。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have gained popularity in task planning for long-horizon manipulation tasks. To enhance the validity of LLM-generated plans, visual demonstrations and online videos have been widely employed to guide the planning process. However, for manipulation tasks involving subtle movements but rich contact interactions, visual perception alone may be insufficient for the LLM to fully interpret the demonstration. Additionally, visual data provides limited information on force-related parameters and conditions, which are crucial for effective execution on real robots. In this paper, we introduce an in-context learning framework that incorporates tactile and force-torque information from human demonstrations to enhance LLMs' ability to generate plans for new task scenarios. We propose a bootstrapped reasoning pipeline that sequentially integrates each modality into a comprehensive task plan. This task plan is then used as a reference for planning in new task configurations. Real-world experiments on two different sequential manipulation tasks demonstrate the effectiveness of our framework in improving LLMs' understanding of multi-modal demonstrations and enhancing the overall planning performance.

大模型多模态机器人操作触觉感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。