用Transformer和大模型让机器人仅靠摄像头完成任务,无需激光雷达等设备。
MissionGPT: Mission Planner for Mobile Robot based on Robotics Transformer Model
- 基于Transformer与大语言模型构建任务规划器,仅依赖摄像头输入。
- 在基础动作上实现超过50%的成功率,无需感知算法支持。
- 适合仓库物流等场景,可推广至各类机器人及多机协同系统。
本文提出一种基于神经网络与Transformer架构的新型任务规划方法,结合大型语言模型(LLMs)。该方法展示了仅通过摄像头数据即可向移动机器人下达任务并成功执行的可能性,无需依赖感知算法。实验中,针对移动机器人的基本动作任务,取得了超过50%的成功率。该方法在仓储物流领域具有实际应用价值,未来有望替代标记、激光雷达、信标等空间定位工具。结论表明,该方法具备可扩展性,适用于任意类型机器人及多机器人系统。
原文摘要 · Abstract (English)
This paper presents a novel approach to building mission planners based on neural networks with Transformer architecture and Large Language Models (LLMs). This approach demonstrates the possibility of setting a task for a mobile robot and its successful execution without the use of perception algorithms, based only on the data coming from the camera. In this work, a success rate of more than 50\% was obtained for one of the basic actions for mobile robots. The proposed approach is of practical importance in the field of warehouse logistics robots, as in the future it may allow to eliminate the use of markings, LiDARs, beacons and other tools for robot orientation in space. In conclusion, this approach can be scaled for any type of robot and for any number of robots.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。