arXiv:2412.00103cs.ROcs.AI2024-12被引 5

用多模态大模型实现零样本找人,机器人能实时应对人员行程变化。

MLLM-Search: A Zero-Shot Approach to Finding People using Multimodal Large Language Models

  • 通过视觉提示生成带空间语义的导航点图,让机器人理解环境布局。
  • 在动态场景中搜索效率优于现有方法,3D仿真测试成功率超90%。
  • 无需训练即可适应新环境,适合医疗等复杂场景的移动机器人使用。

在以人为中心的环境中(如医疗场所)进行自主寻人极具挑战性,因机器人缺乏对人员日程、计划或位置的先验知识,且需应对实时事件引发的计划变更。本文提出MLLM-Search,一种基于多模态大语言模型(MLLM)的零样本寻人架构,用于解决移动机器人在事件驱动场景下搜寻人员的问题。该方法引入新颖的视觉提示策略,生成具有空间锚定性的路径点地图,以拓扑图表示可通行路径点,以语义标签表示区域。结合区域规划器(依据语义相关性选择下一搜索区域)与路径规划器(通过独特的空间思维链提示生成路径,考虑语义相关物体及局部空间上下文)。在多种环境的3D逼真仿真中验证了其性能,结果显示在人员日程变动条件下搜索成功率超过90%。消融实验验证了核心设计的有效性,与前沿方法对比表明,在搜索效率上显著领先。真实世界实验中,移动机器人在多房间建筑中成功泛化至未见环境完成寻人任务。

原文摘要 · Abstract (English)

Robotic search of people in human-centered environments, including healthcare settings, is challenging as autonomous robots need to locate people without complete or any prior knowledge of their schedules, plans or locations. Furthermore, robots need to be able to adapt to real-time events that can influence a person's plan in an environment. In this paper, we present MLLM-Search, a novel zero-shot person search architecture that leverages multimodal large language models (MLLM) to address the mobile robot problem of searching for a person under event-driven scenarios with varying user schedules. Our approach introduces a novel visual prompting method to provide robots with spatial understanding of the environment by generating a spatially grounded waypoint map, representing navigable waypoints by a topological graph and regions by semantic labels. This is incorporated into a MLLM with a region planner that selects the next search region based on the semantic relevance to the search scenario, and a waypoint planner which generates a search path by considering the semantically relevant objects and the local spatial context through our unique spatial chain-of-thought prompting approach. Extensive 3D photorealistic experiments were conducted to validate the performance of MLLM-Search in searching for a person with a changing schedule in different environments. An ablation study was also conducted to validate the main design choices of MLLM-Search. Furthermore, a comparison study with state-of-the art search methods demonstrated that MLLM-Search outperforms existing methods with respect to search efficiency. Real-world experiments with a mobile robot in a multi-room floor of a building showed that MLLM-Search was able to generalize to finding a person in a new unseen environment.

机器人寻人多模态大模型零样本空间推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。