arXiv:2502.08903cs.ROcs.AI2025-02被引 7

让机器人通过2D图像自动获取3D空间信息,提升任务成功率至96%

3D-Grounded Vision-Language Framework for Robotic Task Planning: Automated Prompt Synthesis and Supervised Reasoning

  • 用2D图像生成3D点云提示,无需人工干预
  • 小语言模型监督输出,使任务成功率达96%
  • 适合需要高精度定位的机器人任务规划

视觉语言模型(VLMs)在场景理解与感知任务中表现优异,使机器人能在动态环境中自适应地规划与执行动作。然而,大多数多模态大模型缺乏稳健的3D场景定位能力,限制了其在精细操作中的应用。此外,识别准确率低、效率差、泛化性弱和可靠性不足等问题也阻碍了其在精密任务中的使用。为此,我们提出一种新框架,包含2D提示生成模块(将2D图像映射到点云)和小型语言模型(SLM)用于监督VLM输出。该模块使训练于2D图像与文本的VLM能自主提取精确3D空间信息,显著增强3D场景理解能力。同时,SLM对VLM输出进行监督,减少幻觉,确保生成可执行的机器人控制代码。本框架无需在新环境中重新训练,提升了成本效益与运行鲁棒性。实验结果表明,所提框架任务成功率达96.0%,优于其他方法。消融实验显示,两个模块均至关重要:移除任一模块后任务成功率下降67%。这些结果验证了框架在提升3D识别、任务规划与机器人执行方面的有效性。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have achieved remarkable success in scene understanding and perception tasks, enabling robots to plan and execute actions adaptively in dynamic environments. However, most multimodal large language models lack robust 3D scene localization capabilities, limiting their effectiveness in fine-grained robotic operations. Additionally, challenges such as low recognition accuracy, inefficiency, poor transferability, and reliability hinder their use in precision tasks. To address these limitations, we propose a novel framework that integrates a 2D prompt synthesis module by mapping 2D images to point clouds, and incorporates a small language model (SLM) for supervising VLM outputs. The 2D prompt synthesis module enables VLMs, trained on 2D images and text, to autonomously extract precise 3D spatial information without manual intervention, significantly enhancing 3D scene understanding. Meanwhile, the SLM supervises VLM outputs, mitigating hallucinations and ensuring reliable, executable robotic control code generation. Our framework eliminates the need for retraining in new environments, thereby improving cost efficiency and operational robustness. Experimental results that the proposed framework achieved a 96.0\% Task Success Rate (TSR), outperforming other methods. Ablation studies demonstrated the critical role of both the 2D prompt synthesis module and the output supervision module (which, when removed, caused a 67\% TSR drop). These findings validate the framework's effectiveness in improving 3D recognition, task planning, and robotic task execution.

机器人规划3D定位视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。