arXiv:2602.04231cs.RO2026-02

用视觉语言统一表征+深度几何先验,提升机器人在复杂场景下的语言抓取能力

GeoLanG: Geometry-Aware Language-Guided Grasping with Unified RGB-D Multimodal Learning

  • 基于CLIP架构的端到端多任务框架,融合视图与语言信息
  • 在OCID-VLG数据集上实现91.3%抓取成功率,真实场景表现显著优于基线
  • 适合需高鲁棒性多模态交互的智能机器人研发人员

语言引导抓取已成为机器人通过自然语言指令识别并操作目标物体的有前景范式,但在杂乱或遮挡场景中仍具挑战。现有方法多采用分阶段流水线,分离物体感知与抓取,导致跨模态融合有限、计算冗余,且在杂乱、遮挡或低纹理场景中泛化能力差。为此,我们提出GeoLanG,一个基于CLIP架构的端到端多任务框架,将视觉与语言输入统一至共享表示空间,实现鲁棒语义对齐与更强泛化能力。为增强遮挡和低纹理条件下的目标区分能力,我们引入深度引导几何模块(DGGM),将深度信息转化为显式几何先验,并注入注意力机制,无额外计算开销。此外,提出自适应密集通道融合,自适应平衡多层特征贡献,生成更具判别性和泛化性的视觉表征。在OCID-VLG数据集及仿真与真实硬件上的大量实验表明,GeoLanG可在复杂杂乱环境中实现精确可靠的语言引导抓取,为面向人机共存场景的多模态机器人操作提供新路径。

原文摘要 · Abstract (English)

Language-guided grasping has emerged as a promising paradigm for enabling robots to identify and manipulate target objects through natural language instructions, yet it remains highly challenging in cluttered or occluded scenes. Existing methods often rely on multi-stage pipelines that separate object perception and grasping, which leads to limited cross-modal fusion, redundant computation, and poor generalization in cluttered, occluded, or low-texture scenes. To address these limitations, we propose GeoLanG, an end-to-end multi-task framework built upon the CLIP architecture that unifies visual and linguistic inputs into a shared representation space for robust semantic alignment and improved generalization. To enhance target discrimination under occlusion and low-texture conditions, we explore a more effective use of depth information through the Depth-guided Geometric Module (DGGM), which converts depth into explicit geometric priors and injects them into the attention mechanism without additional computational overhead. In addition, we propose Adaptive Dense Channel Integration, which adaptively balances the contributions of multi-layer features to produce more discriminative and generalizable visual representations. Extensive experiments on the OCID-VLG dataset, as well as in both simulation and real-world hardware, demonstrate that GeoLanG enables precise and robust language-guided grasping in complex, cluttered environments, paving the way toward more reliable multimodal robotic manipulation in real-world human-centric settings.

机器人抓取多模态学习语言引导几何先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。