arXiv:2501.05014cs.ROcs.AI2025-01被引 57

用自然语言生成无人机大规模飞行任务,提升规划效率与精度。

UAV-VLA: Vision-Language-Action System for Large Scale Aerial Mission Generation

  • 结合卫星图像与大模型,通过文本指令生成飞行路径与动作
  • 轨迹长度差异减少22%,目标定位误差降低至34.22米(欧氏距离)
  • 适合非专业用户快速部署复杂航拍任务

UAV-VLA(视觉-语言-动作)系统是一种用于与空基机器人交互的工具。通过整合卫星影像处理、视觉语言模型(VLM)及GPT的强大能力,该系统使用户仅需简单文本请求即可生成通用飞行路径与行动方案。系统利用卫星图像提供的丰富上下文信息,提升决策与任务规划能力。VLM的视觉分析与GPT的语言理解相结合,可输出完整路径与动作集,使空中作业更高效、易用。新方法在轨迹长度差异上改善22%,在K近邻(KNN)方法中目标定位的均方误差降低至34.22米(欧氏距离)。

原文摘要 · Abstract (English)

The UAV-VLA (Visual-Language-Action) system is a tool designed to facilitate communication with aerial robots. By integrating satellite imagery processing with the Visual Language Model (VLM) and the powerful capabilities of GPT, UAV-VLA enables users to generate general flight paths-and-action plans through simple text requests. This system leverages the rich contextual information provided by satellite images, allowing for enhanced decision-making and mission planning. The combination of visual analysis by VLM and natural language processing by GPT can provide the user with the path-and-action set, making aerial operations more efficient and accessible. The newly developed method showed the difference in the length of the created trajectory in 22% and the mean error in finding the objects of interest on a map in 34.22 m by Euclidean distance in the K-Nearest Neighbors (KNN) approach.

无人机多模态任务生成自然语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。