arXiv:2503.19510cs.ROcs.AI2025-03被引 17

融合深度与可见光信息,让机器人更懂语言指令地完成复杂操作。

RoboFlamingo-Plus: Fusion of Depth and RGB Perception with Vision-Language Models for Enhanced Robotic Manipulation

  • 用预训练视觉变换器加重采样技术融合深度与图像数据。
  • 在挑战性场景中任务成功率提升10%-20%。
  • 适合研究多模态机器人操控或视觉语言模型应用者。

随着机器人技术向更复杂的多模态交互与操作任务演进,先进视觉语言模型(VLMs)的集成已成为关键驱动力。尽管现有方法取得进展,但在3D环境中融合深度与RGB信息并执行语言指导任务仍面临挑战。为此,我们改进了RoboFlamingo框架,提出RoboFlamingo-Plus,将深度数据融入VLM以显著提升机器人操作性能。通过将预训练视觉变换器(ViT)与重采样技术结合,实现RGB与深度信息的精细融合,并利用交叉注意力机制对齐语义线索,增强多模态理解。其创新在于针对深度数据的输入适配、基于预训练重采样器的特征提取及最优特征融合策略。实验表明,RoboFlamingo-Plus相较当前方法在机器人操作任务上提升10%-20%,显著推进该领域发展。代码与模型权重已公开于RoboFlamingo-Plus。

原文摘要 · Abstract (English)

As robotic technologies advancing towards more complex multimodal interactions and manipulation tasks, the integration of advanced Vision-Language Models (VLMs) has become a key driver in the field. Despite progress with current methods, challenges persist in fusing depth and RGB information within 3D environments and executing tasks guided by linguistic instructions. In response to these challenges, we have enhanced the existing RoboFlamingo framework by introducing RoboFlamingo-Plus, which incorporates depth data into VLMs to significantly improve robotic manipulation performance. Our research achieves a nuanced fusion of RGB and depth information by integrating a pre-trained Vision Transformer (ViT) with a resampling technique, closely aligning this combined data with linguistic cues for superior multimodal understanding. The novelty of RoboFlamingo-Plus lies in its adaptation of inputs for depth data processing, leveraging a pre-trained resampler for depth feature extraction, and employing cross-attention mechanisms for optimal feature integration. These improvements allow RoboFlamingo-Plus to not only deeply understand 3D environments but also easily perform complex, language-guided tasks in challenging settings. Experimental results show that RoboFlamingo-Plus boosts robotic manipulation by 10-20% over current methods, marking a significant advancement. Codes and model weights are public at RoboFlamingo-Plus.

机器人操控视觉语言模型多模态融合深度感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。