用Mamba模型实现多尺度融合,让机器人更快更准地听懂指令抓东西。
GraspMamba: A Mamba-based Language-driven Grasp Detection Framework with Hierarchical Feature Learning
- 基于Mamba架构分层融合视觉与语言特征
- 在复杂场景下仍保持高精度与快速推理
- 适合需要实时响应的工业机器人应用
抓取检测是机器人完成众多工业任务的基础。然而,现有语言驱动模型在处理杂乱图像、长文本描述或推理速度慢方面表现不佳。我们提出GraspMamba,一种基于Mamba视觉的新型语言驱动抓取检测方法,通过分层特征融合有效应对这些挑战。该方法首次在多尺度上利用Mamba骨干网络提取视觉与语言特征,显著增强跨模态融合能力。实验表明,GraspMamba在多个基准测试中明显优于现有方法,且在真实机器人实验中验证了其快速推理优势。
原文摘要 · Abstract (English)
Grasp detection is a fundamental robotic task critical to the success of many industrial applications. However, current language-driven models for this task often struggle with cluttered images, lengthy textual descriptions, or slow inference speed. We introduce GraspMamba, a new language-driven grasp detection method that employs hierarchical feature fusion with Mamba vision to tackle these challenges. By leveraging rich visual features of the Mamba-based backbone alongside textual information, our approach effectively enhances the fusion of multimodal features. GraspMamba represents the first Mamba-based grasp detection model to extract vision and language features at multiple scales, delivering robust performance and rapid inference time. Intensive experiments show that GraspMamba outperforms recent methods by a clear margin. We validate our approach through real-world robotic experiments, highlighting its fast inference speed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。