用大模型实时分析语义与空间信息,让无人机搜物更快更准。
Spatial-Semantic Reasoning using Large Language Models for Efficient UAV Search Operations

- 大模型解析自然语言指令,结合物体和空间信息判断搜索优先级。
- 实测任务时长缩短,真实与仿真环境下搜寻准确率均超90%。
- 适合需要快速响应的无人机搜救、巡检等实时场景。
我们提出一种面向无人飞行器(UAV)的实时语义导航框架,旨在提升目标导航(ObjectNav)任务的时间效率。核心是利用大型语言模型(LLM)理解用户提供的自然语言指令,并对检测到的物体及空间上下文进行语义推理,以优先规划高概率搜索区域。系统融合实时目标检测、3D空间建图以及多项式样条插值,实现平滑且可行的无人机轨迹规划。与依赖离线推理或受仿真环境限制动作空间的方法不同,本框架可实时运行,并根据新观测持续更新语义相关性。在模拟与真实环境中的实验表明,任务时长显著减少,同时保持高搜寻准确率,验证了基于大模型的语义推理在高效无人机目标导航中的有效性。
原文摘要 · Abstract (English)
We present a real-time semantic navigation framework for Unmanned Aerial Vehicles (UAVs) focused on improving time efficiency in the Object Goal Navigation (ObjectNav) task. Central to our approach is a Large Language Model (LLM) that interprets user-provided natural language instructions and performs semantic reasoning over detected objects and spatial context to prioritize high-probability search regions. The system combines real-time object detection, 3D spatial mapping, and polynomial spline interpolation for smooth and feasible UAV trajectory planning. Unlike prior methods that rely on offline reasoning or simulator-constrained action spaces, our framework can operate in real time, continuously updating semantic relevance based on new observations. Experiments in both simulated and real-world settings demonstrate reductions in mission duration while maintaining high search accuracy, underscoring the effectiveness of LLM-guided reasoning for time- efficient UAV-based ObjectNav.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。