用大模型+视觉语言模型实现无人机与地面机器人协同语义导航与操作
Hierarchical Language Models for Semantic Navigation and Manipulation in an Aerial-Ground Robotic System
- 分层架构:大模型负责任务分解与全局地图构建,视觉语言模型实现精准物体定位
- 引入GridMask提升空间感知精度,在动态环境中实现零样本适应与可靠操作
- 首次实现异构无人机-地面机器人系统在长时序任务中的语义协作,适合复杂场景应用
异构多机器人系统在需要协调混合协作的复杂任务中具有巨大潜力。然而,现有依赖静态或特定任务模型的方法通常缺乏在多样化任务和动态环境中的泛化能力,凸显出将高层推理与底层执行跨异构智能体连接的通用智能需求。为此,我们提出一种分层多模态框架,融合提示式大语言模型(LLM)与微调的视觉语言模型(VLM)。系统层面,LLM执行分层任务分解并构建全局语义地图;VLM提供语义感知与物体定位,其中提出的GridMask显著提升VLM的空间精度,确保精细操作的可靠性。无人机利用该全局地图生成语义路径,并引导地面机器人的局部导航与操作,即使在目标缺失或场景模糊的情况下也能保持稳健协同。通过长时间序列物体排列任务的大量仿真与真实世界实验验证,框架展现出零样本适应性、鲁棒的语义导航及动态环境中的可靠操作能力。据我们所知,这是首个将基于VLM的感知与基于LLM的推理结合,用于异构空中-地面机器人系统全局高层次任务规划与执行的工作。
原文摘要 · Abstract (English)
Heterogeneous multirobot systems show great potential in complex tasks requiring coordinated hybrid cooperation. However, existing methods that rely on static or task-specific models often lack generalizability across diverse tasks and dynamic environments. This highlights the need for generalizable intelligence that can bridge high-level reasoning with low-level execution across heterogeneous agents. To address this, we propose a hierarchical multimodal framework that integrates a prompted large language model (LLM) with a fine-tuned vision-language model (VLM). At the system level, the LLM performs hierarchical task decomposition and constructs a global semantic map, while the VLM provides semantic perception and object localization, where the proposed GridMask significantly enhances the VLM's spatial accuracy for reliable fine-grained manipulation. The aerial robot leverages this global map to generate semantic paths and guide the ground robot's local navigation and manipulation, ensuring robust coordination even in target-absent or ambiguous scenarios. We validate the framework through extensive simulation and real-world experiments on long-horizon object arrangement tasks, demonstrating zero-shot adaptability, robust semantic navigation, and reliable manipulation in dynamic environments. To the best of our knowledge, this work presents the first heterogeneous aerial-ground robotic system that integrates VLM-based perception with LLM-driven reasoning for global high-level task planning and execution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。