用大模型推理能力精准生成抓取姿态,让机器人更懂上下文
RT-Grasp: Reasoning Tuning Robotic Grasping via Multi-modal Large Language Model
- 训练时加入推理阶段,让大模型学会生成数值化抓取动作
- 在真实场景中抓取成功率显著提升,支持对话式调整
- 适合想用大模型做机器人控制的开发者和研究者
大型语言模型(LLMs)虽展现出强大推理能力,但在机器人领域仍多用于文本规划,受限于其输出形式。本文提出推理调优(Reasoning Tuning)方法,在训练中引入推理阶段,利用多模态大模型的先验知识与推理能力,生成上下文感知、可对话调整的数值化抓取姿态。为此构建了专门的数据集 Reasoning Tuning VLM Grasp,支持大模型适配机器人抓取任务。在多个抓取数据集及真实实验中验证,多模态大模型可有效完成数值预测任务,拓展其在机器人控制中的应用,弥合文本规划与直接控制之间的鸿沟。
原文摘要 · Abstract (English)
Recent advances in Large Language Models (LLMs) have showcased their remarkable reasoning capabilities, making them influential across various fields. However, in robotics, their use has primarily been limited to manipulation planning tasks due to their inherent textual output. This paper addresses this limitation by investigating the potential of adopting the reasoning ability of LLMs for generating numerical predictions in robotics tasks, specifically for robotic grasping. We propose Reasoning Tuning, a novel method that integrates a reasoning phase before prediction during training, leveraging the extensive prior knowledge and advanced reasoning abilities of LLMs. This approach enables LLMs, notably with multi-modal capabilities, to generate accurate numerical outputs like grasp poses that are context-aware and adaptable through conversations. Additionally, we present the Reasoning Tuning VLM Grasp dataset, carefully curated to facilitate the adaptation of LLMs to robotic grasping. Extensive validation on both grasping datasets and real-world experiments underscores the adaptability of multi-modal LLMs for numerical prediction tasks in robotics. This not only expands their applicability but also bridges the gap between text-based planning and direct robot control, thereby maximizing the potential of LLMs in robotics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。