用大模型当无人机飞行员,能听懂指令自动飞。
VLN-Pilot: Large Vision-Language Model as an Autonomous Indoor Drone Operator
- 用大视觉语言模型理解自然语言指令并结合视觉感知规划飞行路径
- 在无GPS室内环境完成复杂多目标长程导航,成功率高
- 适合需要人机交互的室内巡检、搜救等场景
本文提出VLN-Pilot,一种由大视觉-语言模型(VLLM)担任室内无人机自主操作员的新框架。通过利用VLLM的多模态推理能力,该框架将自由形式的自然语言指令与视觉观察对齐,在无GPS的室内环境中规划并执行飞行轨迹。相比传统规则或几何路径规划方法,本框架融合语言驱动的语义理解与视觉感知,实现上下文感知的高层飞行行为,几乎无需任务定制化工程。VLN-Pilot支持无人机完全自主地遵循指令,具备空间关系推理、障碍物规避及对突发情况的动态响应能力。我们在自建的逼真室内仿真基准上验证了该框架,结果表明基于VLLM的智能体在包含多个语义目标的复杂指令跟随任务中表现出高成功率。实验表明,用语言引导的自主代理替代远程操控员,可显著降低操作负担,提升安全性与任务灵活性,为室内无人机在巡检、搜救和设施监控中的规模化应用提供新路径。
原文摘要 · Abstract (English)
This paper introduces VLN-Pilot, a novel framework in which a large Vision-and-Language Model (VLLM) assumes the role of a human pilot for indoor drone navigation. By leveraging the multimodal reasoning abilities of VLLMs, VLN-Pilot interprets free-form natural language instructions and grounds them in visual observations to plan and execute drone trajectories in GPS-denied indoor environments. Unlike traditional rule-based or geometric path-planning approaches, our framework integrates language-driven semantic understanding with visual perception, enabling context-aware, high-level flight behaviors with minimal task-specific engineering. VLN-Pilot supports fully autonomous instruction-following for drones by reasoning about spatial relationships, obstacle avoidance, and dynamic reactivity to unforeseen events. We validate our framework on a custom photorealistic indoor simulation benchmark and demonstrate the ability of the VLLM-driven agent to achieve high success rates on complex instruction-following tasks, including long-horizon navigation with multiple semantic targets. Experimental results highlight the promise of replacing remote drone pilots with a language-guided autonomous agent, opening avenues for scalable, human-friendly control of indoor UAVs in tasks such as inspection, search-and-rescue, and facility monitoring. Our results suggest that VLLM-based pilots may dramatically reduce operator workload while improving safety and mission flexibility in constrained indoor environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。