让机器人通过自然语言理解并自动定位施工任务点
Task-Aware Positioning for Improvisational Tasks in Mobile Construction Robots via an AI Agent with Multi-LMM Modules
- 分三模块并行处理任务理解、图纸导航与视觉定位
- 在三项测试中实现92.2%的任务点识别与定位成功率
- 适合需要灵活应对非预设任务的建筑机器人研究者
由于施工现场环境持续变化,许多任务以临时应变方式出现。现有移动施工机器人研究难以应对此类即兴任务,因任务位置、发生时机及执行所需上下文信息无法提前获知。本文提出一种基于多模态大模型(LMM)模块的智能体,能理解自然语言描述的即兴任务,识别任务所需位置并自主定位。该智能体将功能分解为三个并行运行的大型多模态模型模块,分别负责任务解析与拆解、基于施工图纸的导航以及视觉推理以发现未预定义的任务位置。实验采用四足机器人实现,针对即兴任务设计的三项测试中,成功率达92.2%,验证了移动施工机器人自主完成非预设任务的可行性。
原文摘要 · Abstract (English)
Due to the ever-changing nature of construction, many tasks on sites occur in an improvisational manner. Existing mobile construction robot studies remain limited in addressing improvisational tasks, where task-required locations, timing of task occurrence, and contextual information required for task execution are not known in advance. We propose an agent that understands improvisational tasks given in natural language, identifies the task-required location, and positions itself. The agent's functionality was decomposed into three Large Multimodal Model (LMM) modules operating in parallel, enabling the application of LMMs for task interpretation and breakdown, construction drawing-based navigation, and visual reasoning to identify non-predefined task-required locations. The agent was implemented with a quadruped robot and achieved a 92.2% success rate for identifying and positioning at task-required locations across three tests designed to assess improvisational task handling. This study enables mobile construction robots to perform non-predefined tasks autonomously.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。