arXiv:2512.21065cs.ROcs.CV2025-12被引 1

用语言指导机器人抓取,提升复杂场景下的精准度与泛化能力。

Language-Guided Grasp Detection with Coarse-to-Fine Learning for Robotic Manipulation

  • 分步融合语言与视觉信息,逐步增强语义对齐。
  • 在两个数据集上超越现有方法,对未见物体和多样指令均表现优异。
  • 适合需要语言控制的机器人抓取任务,实机部署验证有效。

抓取是机器人操作中最基础也最具挑战性的能力之一,尤其在非结构化、杂乱且语义多样的环境中。近年来,语言引导的操作研究日益增多,机器人不仅需感知场景,还需理解任务相关的自然语言指令。然而,现有语言条件抓取方法多依赖浅层融合策略,导致语义定位不足,语言意图与视觉抓取推理之间对齐较弱。本文提出语言引导抓取检测(LGGD),采用粗到细学习范式,利用基于CLIP的视觉与文本嵌入,在分层跨模态融合流程中逐步注入语言线索,促进视觉特征重建过程中的细粒度语义对齐,提升预测抓取与任务指令的契合度。此外,引入语言条件动态卷积头(LDCH),根据句级特征混合多个卷积专家,实现指令自适应的粗粒度掩码与抓取预测;最终细化模块进一步增强复杂场景下的抓取一致性与鲁棒性。在OCID-VLG和Grasp-Anything++数据集上的实验表明,LGGD优于现有方法,对未见物体及多样化语言查询具备强泛化能力。真实机器人平台部署验证了该方法在执行精确、指令驱动抓取动作方面的实用性。代码将在论文录用后公开。

原文摘要 · Abstract (English)

Grasping is one of the most fundamental challenging capabilities in robotic manipulation, especially in unstructured, cluttered, and semantically diverse environments. Recent researches have increasingly explored language-guided manipulation, where robots not only perceive the scene but also interpret task-relevant natural language instructions. However, existing language-conditioned grasping methods typically rely on shallow fusion strategies, leading to limited semantic grounding and weak alignment between linguistic intent and visual grasp reasoning.In this work, we propose Language-Guided Grasp Detection (LGGD) with a coarse-to-fine learning paradigm for robotic manipulation. LGGD leverages CLIP-based visual and textual embeddings within a hierarchical cross-modal fusion pipeline, progressively injecting linguistic cues into the visual feature reconstruction process. This design enables fine-grained visual-semantic alignment and improves the feasibility of the predicted grasps with respect to task instructions. In addition, we introduce a language-conditioned dynamic convolution head (LDCH) that mixes multiple convolution experts based on sentence-level features, enabling instruction-adaptive coarse mask and grasp predictions. A final refinement module further enhances grasp consistency and robustness in complex scenes.Experiments on the OCID-VLG and Grasp-Anything++ datasets show that LGGD surpasses existing language-guided grasping methods, exhibiting strong generalization to unseen objects and diverse language queries. Moreover, deployment on a real robotic platform demonstrates the practical effectiveness of our approach in executing accurate, instruction-conditioned grasp actions. The code will be released publicly upon acceptance.

机器人抓取语言引导跨模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。