让异构机器人在未知环境里自主协作,靠大模型理解指令并分工执行。
D-VLC: Decentralized Vision-Language Collaboration for Heterogeneous Embodied Multi-Robot Systems in Unknown Environments

- 去中心化异步推理+轻量信息共享,支持不同机器人协同。
- 在多种场景下任务成功率超70%,耗时比基线少55.8%。
- 无需针对特定任务或机器人训练,适合复杂未知环境应用。
异构机器人集群通过并行协作与能力互补,可提升复杂任务的执行效率。然而传统基于规则的方法依赖预设任务模型和专用决策程序,难以理解复杂语义指令并协调异构机器人。大语言模型(LLM)具备强大的语言理解和任务推理能力,使多机器人系统能解析指令、分解任务并按语义分配角色。视觉语言模型(VLM)进一步引入视觉感知,使机器人能推理物理环境中物体、区域及空间关系。但现有基于LLM/VLM的方法通常依赖已知地图和集中式同步决策,限制了其在异构机器人和未见任务上的泛化能力。为此,我们提出一个框架,融合去中心化异步推理、轻量信息共享、能力感知协作与统一动作接口,使通用VLM能够生成无需任务或机器人特训的机器人专属动作。在多样场景和多个VLM上的实验表明,任务成功率超过70%,完成时间相较几何贪婪基线减少最高55.8%。
原文摘要 · Abstract (English)
Multi-robot systems, particularly heterogeneous robot swarms, can improve the efficiency of complex task execution through parallel collaboration and complementary capabilities. However, conventional rule-based methods rely on predefined task models and specialized decision making programs, making it difficult to understand complex semantic instructions and coordinate heterogeneous robots. LLMs introduce strong language understanding and task reasoning capabilities, allowing multi-robot systems to interpret instructions, decompose tasks, and assign roles according to task semantics. VLMs further incorporate visual perception, enabling robots to reason about objects, regions, and spatial relationships in physical environments. Nevertheless, existing LLM/VLM based methods often depend on known maps, centralized and synchronized decision making, limiting their generalization to heterogeneous robots and unseen tasks. We therefore propose a framework that combines decentralized asynchronous reasoning, lightweight information sharing, capability aware collaboration, and a unified action interface, enabling general purpose VLMs to generate robot specific actions executed by learning free experts without task or robot specific training. Experiments across diverse scenarios and multiple VLMs show success rates above 70\%, with completion time reduced by up to 55.8\% relative to the geometric greedy baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。