arXiv:2411.12520cs.ROcs.CV2024-11

VMGNet用状态空间模型实现低算力高精度抓取,适合实时机器人应用。

VMGNet: A Low Computational Complexity Robotic Grasping Network Based on VMamba with Multi-Scale Feature Fusion

  • 引入视觉状态空间模型,实现线性计算复杂度,大幅降低算力消耗。
  • 设计轻量级多尺度融合模块,提升抓取准确率,在两个公开数据集上达顶尖水平。
  • 真实场景测试成功率达94.4%,适合工业级实时抓取任务。

尽管基于深度学习的机器人抓取技术展现出强大适应性,但其计算复杂度显著增加,难以满足高实时性场景需求。为此,我们提出一种低计算复杂度且高精度的抓取模型VMGNet。首次将视觉状态空间(Visual State Space)引入抓取领域,实现线性计算复杂度,极大降低模型计算开销。同时,为提升模型精度,提出一种高效轻量的多尺度特征融合模块——Fusion Bridge Module,用于提取并融合不同尺度信息。此外,设计新的损失函数计算方法,增强子任务间重要性差异,提升模型拟合能力。实验表明,VMGNet在我们的设备上仅需8.7G浮点运算,推理时间仅为8.1ms。在Cornell和Jacquard公开数据集上达到当前最优性能。为验证实际应用效果,我们在多物体场景中进行真实抓取实验,VMGNet在现实任务中取得94.4%的成功率。真实机器人抓取视频见:https://youtu.be/S-QHBtbmLc4。

原文摘要 · Abstract (English)

While deep learning-based robotic grasping technology has demonstrated strong adaptability, its computational complexity has also significantly increased, making it unsuitable for scenarios with high real-time requirements. Therefore, we propose a low computational complexity and high accuracy model named VMGNet for robotic grasping. For the first time, we introduce the Visual State Space into the robotic grasping field to achieve linear computational complexity, thereby greatly reducing the model's computational cost. Meanwhile, to improve the accuracy of the model, we propose an efficient and lightweight multi-scale feature fusion module, named Fusion Bridge Module, to extract and fuse information at different scales. We also present a new loss function calculation method to enhance the importance differences between subtasks, improving the model's fitting ability. Experiments show that VMGNet has only 8.7G Floating Point Operations and an inference time of 8.1 ms on our devices. VMGNet also achieved state-of-the-art performance on the Cornell and Jacquard public datasets. To validate VMGNet's effectiveness in practical applications, we conducted real grasping experiments in multi-object scenarios, and VMGNet achieved an excellent performance with a 94.4% success rate in real-world grasping tasks. The video for the real-world robotic grasping experiments is available at https://youtu.be/S-QHBtbmLc4.

机器人抓取状态空间低算力多尺度融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。