arXiv:2510.02104cs.RO2025-10被引 1

让机器人理解模糊指令,精准抓取物体的细微部位

LangGrasp: Leveraging Fine-Tuned LLMs for Language Interactive Robot Grasping with Ambiguous Instructions

  • 用微调大模型解析语言中的隐含意图
  • 实现从物体级到部件级的高精度抓取
  • 适合复杂场景下需要理解指令的机器人任务

现有语言驱动抓取方法难以充分处理包含隐含意图的模糊指令。为此,我们提出 LangGrasp,一种新型语言交互式机器人抓取框架。该框架融合微调的大语言模型(LLMs),利用其强大的常识理解与环境感知能力,从语言指令中推断隐含意图,并明确任务需求及目标操作对象。此外,设计的点云定位模块基于2D部件分割,实现场景中部分点云的定位,使抓取操作从粗粒度物体级拓展至细粒度部件级。实验结果表明,LangGrasp能准确解析模糊指令中的隐含意图,识别出未明说但对任务完成至关重要的关键操作与目标信息。同时,通过整合环境信息动态选择最优抓取姿态,实现从物体级到部件级的高精度抓取,显著提升机器人在非结构化环境中的适应性与任务执行效率。更多信息与代码见:https://github.com/wu467/LangGrasp。

原文摘要 · Abstract (English)

The existing language-driven grasping methods struggle to fully handle ambiguous instructions containing implicit intents. To tackle this challenge, we propose LangGrasp, a novel language-interactive robotic grasping framework. The framework integrates fine-tuned large language models (LLMs) to leverage their robust commonsense understanding and environmental perception capabilities, thereby deducing implicit intents from linguistic instructions and clarifying task requirements along with target manipulation objects. Furthermore, our designed point cloud localization module, guided by 2D part segmentation, enables partial point cloud localization in scenes, thereby extending grasping operations from coarse-grained object-level to fine-grained part-level manipulation. Experimental results show that the LangGrasp framework accurately resolves implicit intents in ambiguous instructions, identifying critical operations and target information that are unstated yet essential for task completion. Additionally, it dynamically selects optimal grasping poses by integrating environmental information. This enables high-precision grasping from object-level to part-level manipulation, significantly enhancing the adaptability and task execution efficiency of robots in unstructured environments. More information and code are available here: https://github.com/wu467/LangGrasp.

机器人抓取大模型应用语言理解部件级操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。