用智能无人机实现精准低空配送,突破传统导航局限
LogisticsVLN: Vision-Language Navigation For Low-Altitude Terminal Delivery Based on Agentic UAVs
- 基于多模态大模型构建模块化飞行配送系统
- 在CARLA仿真环境中验证系统可行性并提升定位精度
- 适合研究无人机自主配送与视觉语言模型应用者
智能物流对细粒度末端配送的需求日益增长,亟需基于无人飞行器(UAV)的自主配送系统。然而,现有末程配送研究多依赖地面机器人,而当前无人机视觉-语言导航(VLN)任务主要针对粗粒度、远距离目标,难以满足精确终端交付需求。为此,我们提出LogisticsVLN,一种基于多模态大语言模型(MLLMs)的可扩展空中配送系统,用于自主终端配送。该系统采用模块化流水线,集成轻量级大语言模型(LLMs)与视觉-语言模型(VLMs),完成任务理解、楼层定位、物体检测及动作决策。为支持该领域的研究与评估,我们在CARLA模拟器中构建了视觉-语言配送(VLD)数据集。在VLD数据集上的实验结果验证了LogisticsVLN系统的可行性。此外,我们对系统各模块进行子任务级评估,为基于基础模型的视觉-语言配送系统提升鲁棒性与实际部署能力提供了重要参考。
原文摘要 · Abstract (English)
The growing demand for intelligent logistics, particularly fine-grained terminal delivery, underscores the need for autonomous UAV (Unmanned Aerial Vehicle)-based delivery systems. However, most existing last-mile delivery studies rely on ground robots, while current UAV-based Vision-Language Navigation (VLN) tasks primarily focus on coarse-grained, long-range goals, making them unsuitable for precise terminal delivery. To bridge this gap, we propose LogisticsVLN, a scalable aerial delivery system built on multimodal large language models (MLLMs) for autonomous terminal delivery. LogisticsVLN integrates lightweight Large Language Models (LLMs) and Visual-Language Models (VLMs) in a modular pipeline for request understanding, floor localization, object detection, and action-decision making. To support research and evaluation in this new setting, we construct the Vision-Language Delivery (VLD) dataset within the CARLA simulator. Experimental results on the VLD dataset showcase the feasibility of the LogisticsVLN system. In addition, we conduct subtask-level evaluations of each module of our system, offering valuable insights for improving the robustness and real-world deployment of foundation model-based vision-language delivery systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。