首个面向室内无人机视觉语言导航的基准数据集与模型
IndoorUAV: Benchmarking Vision-Language UAV Navigation in Continuous Indoor Environments
- 构建包含16000条轨迹的室内无人机导航数据集
- 支持长程与短程导航,指令粒度可调
- 适合研究空中机器人、多模态推理的学者
视觉语言导航(VLN)使智能体能根据自然语言指令在复杂环境中导航。尽管现有研究多聚焦于地面机器人或室外无人机,但室内无人机的视觉语言导航仍缺乏探索,而这一方向对巡检、配送、搜救等实际应用具有重要意义。为此,我们提出全新基准 IndoorUAV,首先从 Habitat 模拟器中收集超过1000个结构丰富的3D室内场景,模拟真实无人机飞行动态,生成多样化的3D导航轨迹,并通过数据增强丰富样本。随后设计自动化标注流程,为每条轨迹生成不同粒度的自然语言指令,最终形成超过16000条高质量轨迹,构成专注于长程导航的 IndoorUAV-VLN 子集;为支持短程规划,通过选取语义关键帧并重构指令,生成 IndoorUAV-VLA 子集。最后,我们提出 IndoorUAV-Agent 模型,采用任务分解与多模态推理策略。该工作为室内空中视觉语言智能体的研究提供重要资源。
原文摘要 · Abstract (English)
Vision-Language Navigation (VLN) enables agents to navigate in complex environments by following natural language instructions grounded in visual observations. Although most existing work has focused on ground-based robots or outdoor Unmanned Aerial Vehicles (UAVs), indoor UAV-based VLN remains underexplored, despite its relevance to real-world applications such as inspection, delivery, and search-and-rescue in confined spaces. To bridge this gap, we introduce \textbf{IndoorUAV}, a novel benchmark and method specifically tailored for VLN with indoor UAVs. We begin by curating over 1,000 diverse and structurally rich 3D indoor scenes from the Habitat simulator. Within these environments, we simulate realistic UAV flight dynamics to collect diverse 3D navigation trajectories manually, further enriched through data augmentation techniques. Furthermore, we design an automated annotation pipeline to generate natural language instructions of varying granularity for each trajectory. This process yields over 16,000 high-quality trajectories, comprising the \textbf{IndoorUAV-VLN} subset, which focuses on long-horizon VLN. To support short-horizon planning, we segment long trajectories into sub-trajectories by selecting semantically salient keyframes and regenerating concise instructions, forming the \textbf{IndoorUAV-VLA} subset. Finally, we introduce \textbf{IndoorUAV-Agent}, a novel navigation model designed for our benchmark, leveraging task decomposition and multimodal reasoning. We hope IndoorUAV serves as a valuable resource to advance research on vision-language embodied AI in the indoor aerial navigation domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。