arXiv:2410.03522cs.RO2024-10被引 6

融合Mamba与Transformer,提升机器人在杂乱环境中的抓取精度与适应性。

HMT-Grasp: A Hybrid Mamba-Transformer Approach for Robot Grasping in Cluttered Environments

  • 结合视觉Mamba与并行卷积-注意力块,同时捕捉局部与全局特征。
  • 在Cornell、Jacquard等数据集上性能超越现有方法,真实机器人实验表现优异。
  • 适合需要高鲁棒性抓取的工业与服务机器人场景。

机器人抓取在工业与服务应用中至关重要,无论面对孤立物体、杂乱物品还是堆叠物体。然而,当前基于卷积神经网络(CNN)和视觉变压器(ViTs)的视觉抓取检测方法往往难以适应多样场景,因过分侧重局部或全局特征,忽略互补信息。本文提出一种新型混合Mamba-Transformer方法,通过融合视觉Mamba与并行卷积-变压器模块,有效整合全局与局部信息,显著提升在各类机器人任务中的适应性、精度与灵活性。为确保公平评估,我们在Cornell、Jacquard及OCID-Grasp数据集上进行了广泛实验,涵盖从简单到复杂的多种场景,并开展模拟与真实机器人实验。结果表明,该方法不仅在标准抓取数据集上超越现有最先进水平,且在仿真与真实机器人应用中均表现出色。

原文摘要 · Abstract (English)

Robot grasping, whether handling isolated objects, cluttered items, or stacked objects, plays a critical role in industrial and service applications. However, current visual grasp detection methods based on Convolutional Neural Networks (CNNs) and Vision Transformers (ViTs) often struggle to adapt to diverse scenarios, as they tend to emphasize either local or global features exclusively, neglecting complementary cues. In this paper, we propose a novel hybrid Mamba-Transformer approach to address these challenges. Our method improves robotic visual grasping by effectively capturing both global and local information through the integration of Vision Mamba and parallel convolutional-transformer blocks. This hybrid architecture significantly improves adaptability, precision, and flexibility across various robotic tasks. To ensure a fair evaluation, we conducted extensive experiments on the Cornell, Jacquard, and OCID-Grasp datasets, ranging from simple to complex scenarios. Additionally, we performed both simulated and real-world robotic experiments. The results demonstrate that our method not only surpasses state-of-the-art techniques on standard grasping datasets but also delivers strong performance in both simulation and real-world robot applications.

机器人抓取视觉感知混合模型多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。