让无人机听懂人话,自动导航到羽毛球馆。
"Hi AirStar, Guide Me to the Badminton Court."
- 用大模型做脑,语音手势控制无人机
- 长距离导航+短距离精控,定位准
- 适合普通用户、教学演示、智能拍摄
无人飞行器在障碍物较少的环境中具有高机动性和三维自由度,可快速接近目标并执行地面机器人难以完成的任务,适用于探索、巡检、航拍和日常辅助。本文提出AirStar,一种以无人机为中心的具身智能平台,将大语言模型作为环境理解、上下文推理和任务规划的核心,支持通过语音和手势进行自然交互,无需遥控器,显著扩大用户群体。系统结合地理空间知识驱动的远距离导航与上下文推理支持的精细近距离控制,实现高效准确的视觉-语言导航(VLN)能力。此外,系统还具备跨模态问答、智能录像和目标跟踪等内置功能。其高度可扩展的框架支持新功能无缝集成,为通用指令驱动的智能无人机代理铺平道路。
原文摘要 · Abstract (English)
Unmanned Aerial Vehicles, operating in environments with relatively few obstacles, offer high maneuverability and full three-dimensional mobility. This allows them to rapidly approach objects and perform a wide range of tasks often challenging for ground robots, making them ideal for exploration, inspection, aerial imaging, and everyday assistance. In this paper, we introduce AirStar, a UAV-centric embodied platform that turns a UAV into an intelligent aerial assistant: a large language model acts as the cognitive core for environmental understanding, contextual reasoning, and task planning. AirStar accepts natural interaction through voice commands and gestures, removing the need for a remote controller and significantly broadening its user base. It combines geospatial knowledge-driven long-distance navigation with contextual reasoning for fine-grained short-range control, resulting in an efficient and accurate vision-and-language navigation (VLN) capability.Furthermore, the system also offers built-in capabilities such as cross-modal question answering, intelligent filming, and target tracking. With a highly extensible framework, it supports seamless integration of new functionalities, paving the way toward a general-purpose, instruction-driven intelligent UAV agent. The supplementary PPT is available at \href{https://buaa-colalab.github.io/airstar.github.io}{https://buaa-colalab.github.io/airstar.github.io}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。