arXiv:2509.26324cs.ROcs.AI2025-09被引 1

用视觉语言模型让多机器人高效协作找物,支持自然语言指令。

COMRES-VLM: Coordinated Multi-Robot Exploration and Search using Vision Language Models

  • 结合视觉语言模型与地图分析,动态生成全局一致的导航点分配。
  • 探索速度比现有方法快10.2%,找物效率提升55.7%。
  • 支持人类用自然语言下达语义指令,适合需要人机协作的场景。

自主探索和未知室内环境中的目标物体搜索对多机器人系统仍是挑战。传统方法通常依赖贪婪的前沿分配策略,缺乏机器人间的有效协同。本文提出基于视觉语言模型的协同多机器人探索与搜索框架(COMRES-VLM),利用视觉语言模型(VLM)实现多机器人系统的智能协调,以高效完成探索与目标搜索任务。该框架整合实时前沿簇提取、拓扑骨架分析,并在共享占用地图、机器人状态及可选自然语言先验的基础上进行VLM推理,生成全局一致的航点分配。大规模模拟实验表明,六机器人环境下,COMRES-VLM持续优于先进协调方法(如容量约束车辆路径问题与Voronoi规划器),探索完成速度提升10.2%,目标搜索效率提高55.7%。尤为关键的是,该方法支持自然语言驱动的目标搜索,使操作者可通过高层语义指令引导,而传统算法无法理解此类输入。

原文摘要 · Abstract (English)

Autonomous exploration and object search in unknown indoor environments remain challenging for multi-robot systems (MRS). Traditional approaches often rely on greedy frontier assignment strategies with limited inter-robot coordination. In this work, we present Coordinated Multi-Robot Exploration and Search using Vision Language Models (COMRES-VLM), a novel framework that leverages Vision Language Models (VLMs) for intelligent coordination of MRS tasked with efficient exploration and target object search. COMRES-VLM integrates real-time frontier cluster extraction and topological skeleton analysis with VLM reasoning over shared occupancy maps, robot states, and optional natural language priors, in order to generate globally consistent waypoint assignments. Extensive experiments in large-scale simulated indoor environments with up to six robots demonstrate that COMRES-VLM consistently outperforms state-of-the-art coordination methods, including Capacitated Vehicle Routing Problem (CVRP) and Voronoi-based planners, achieving 10.2\% faster exploration completion and 55.7\% higher object search efficiency. Notably, COMRES-VLM enables natural language-based object search capabilities, allowing human operators to provide high-level semantic guidance that traditional algorithms cannot interpret.

多机器人视觉语言模型协同探索自然语言控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。