arXiv:2604.16298cs.CVcs.RO2026-04中稿 · CVPR被引 1

将无人机导航拆解为精细认知模块,实现零样本精准路径规划。

FineCog-Nav: Integrating Fine-grained Cognitive Modules for Zero-shot Multimodal UAV Navigation

论文配图:FineCog-Nav: Integrating Fine-grained Cognitive Modules for Zero-shot Multimodal UAV Navigation
图 1 · 摘自论文原文
  • 按人类认知分解语言、感知、记忆等模块,各用专用小模型协同工作。
  • 在300条航线的细粒度评测中,指令遵循率和长程规划能力显著提升。
  • 适合研究零样本视觉语言导航与可解释智能体设计的学者参考。

无人机视觉语言导航(VLN)要求智能体从第一人称视角,在复杂三维环境中根据模糊的多步指令完成长期导航任务。现有零样本方法受限于大模型依赖、通用提示和松散模块协作。本文提出 FineCog-Nav,一种受人类认知启发的自上而下框架,将导航细分为语言处理、感知、注意力、记忆、想象、推理和决策等模块。每个模块由中等规模基础模型驱动,配合角色化提示与结构化输入输出协议,实现高效协作与更高可解释性。为支持细粒度评估,构建了 AerialVLN-Fine 基准,包含从 AerialVLN 派生的300条轨迹,具备句子级指令-轨迹对齐及含明确视觉终点和地标参考的优化指令。实验表明,FineCog-Nav 在指令遵循、长程规划与未见环境泛化方面持续优于零样本基线,验证了细粒度认知模块化在零样本空域导航中的有效性。

原文摘要 · Abstract (English)

UAV vision-language navigation (VLN) requires an agent to navigate complex 3D environments from an egocentric perspective while following ambiguous multi-step instructions over long horizons. Existing zero-shot methods remain limited, as they often rely on large base models, generic prompts, and loosely coordinated modules. In this work, we propose FineCog-Nav, a top-down framework inspired by human cognition that organizes navigation into fine-grained modules for language processing, perception, attention, memory, imagination, reasoning, and decision-making. Each module is driven by a moderate-sized foundation model with role-specific prompts and structured input-output protocols, enabling effective collaboration and improved interpretability. To support fine-grained evaluation, we construct AerialVLN-Fine, a curated benchmark of 300 trajectories derived from AerialVLN, with sentence-level instruction-trajectory alignment and refined instructions containing explicit visual endpoints and landmark references. Experiments show that FineCog-Nav consistently outperforms zero-shot baselines in instruction adherence, long-horizon planning, and generalization to unseen environments. These results suggest the effectiveness of fine-grained cognitive modularization for zero-shot aerial navigation. Project page: https://smartdianlab.github.io/projects-FineCogNav.

无人机导航视觉语言导航认知模块零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。