多机器人协作导航新方法,提升探索效率并降低通信开销
Enhancing Multi-Robot Semantic Navigation Through Multimodal Chain-of-Thought Score Collaboration
- 用多模态思维链评估探索价值,结合视觉与语言模型
- 通过全局语义地图减少通信量,稳定输出导航决策
- 适合家庭服务机器人团队协同探索陌生环境
理解人类如何协作利用语义知识探索陌生环境并决定导航方向,对家用多机器人系统至关重要。以往方法多采用单机器人集中式规划,严重限制了探索效率。近期研究虽引入多机器人去中心化规划,但常忽略通信成本。本文提出多模态思维链协同导航(MCoCoNav),一种模块化方法,利用多模态思维链为多机器人规划协同语义导航。MCoCoNav结合视觉感知与视觉语言模型(VLMs),通过概率评分评估探索价值,从而降低时间成本并实现稳定输出。同时,使用全局语义地图作为通信桥梁,在最小化通信开销的同时整合观测结果。在HM3D_v0.2和MP3D数据集上的实验验证了该方法的有效性。代码已开源。
原文摘要 · Abstract (English)
Understanding how humans cooperatively utilize semantic knowledge to explore unfamiliar environments and decide on navigation directions is critical for house service multi-robot systems. Previous methods primarily focused on single-robot centralized planning strategies, which severely limited exploration efficiency. Recent research has considered decentralized planning strategies for multiple robots, assigning separate planning models to each robot, but these approaches often overlook communication costs. In this work, we propose Multimodal Chain-of-Thought Co-Navigation (MCoCoNav), a modular approach that utilizes multimodal Chain-of-Thought to plan collaborative semantic navigation for multiple robots. MCoCoNav combines visual perception with Vision Language Models (VLMs) to evaluate exploration value through probabilistic scoring, thus reducing time costs and achieving stable outputs. Additionally, a global semantic map is used as a communication bridge, minimizing communication overhead while integrating observational results. Guided by scores that reflect exploration trends, robots utilize this map to assess whether to explore new frontier points or revisit history nodes. Experiments on HM3D_v0.2 and MP3D demonstrate the effectiveness of our approach. Our code is available at https://github.com/FrankZxShen/MCoCoNav.git.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。