让无人机在看见目标后精准导航,提升视觉语言定位精度。
See-and-Reach: Precise Vision-Language Navigation for UAVs within the Field of View

- 基于动态3D方向引导,实现高分辨率视觉与空间的精准对齐。
- 在2717条轨迹上实现13.82%成功率提升,显著优于基线模型。
- 适合需要精确末端导航的无人机应用场景,如巡检与救援。
无人机视觉语言导航(UAV-VLN)通常被建模为整体搜索与到达任务,长距离目标发现与最终接近联合优化与评估,难以诊断空中智能体的关键能力:即目标进入视场后能否准确定位并转化为精确3D运动。为此,我们提出UAV-VLN-FOV任务,专门分离“看见并抵达”阶段,支持更诊断性的终端到达评估。进一步提出3DG-VLN框架,通过动态3D方向提示,增强细粒度视觉定位与空间方向对齐。该方法自适应处理高分辨率前视与俯视观测,保留精细视觉与几何细节;在闭环导航中在线更新目标相对方向,减少方向漂移。我们构建了一个专用高分辨率基准,包含2717条轨迹、目标导向的高层指令、高分辨率前后视和俯视视角观测,以及连续3D航点标注。实验表明,3DG-VLN超越竞争基线,在成功率上提升13.82%。真实世界测试验证了其在实际见-达导航中的潜力。代码与数据集开源:https://github.com/xuefanfu/3DG-VLN。
原文摘要 · Abstract (English)
UAV Vision-Language Navigation (UAV-VLN) is typically formulated as a holistic search-and-reach problem, where long-range target discovery and final target approach are optimized and evaluated jointly. This formulation makes it difficult to assess a critical capability of aerial embodied agents, namely whether a UAV can accurately ground a visible target and translate vision-language evidence into precise 3D motion once the target enters its field of view. To address this limitation, we introduce UAV-VLN-FOV, a target-visible navigation task that isolates the see-and-reach stage and enables a more diagnostic evaluation of terminal reaching ability. We further propose 3DG-VLN, a vision-language waypoint prediction framework guided by dynamic 3D direction cues to enhance fine-grained visual grounding and spatial direction alignment for precise target reaching. Specifically, 3DG-VLN adaptively processes high-resolution front-view and downward-view observations to preserve fine-grained visual and geometric details for target grounding. It also updates the target-relative direction online during closed-loop navigation, allowing the agent to maintain spatial alignment with the target and reduce accumulated direction drift. To support this task, we construct a dedicated high-resolution benchmark which contains 2,717 trajectories with target-oriented high-level instructions, high-resolution front-view and downward-view egocentric observations, and continuous 3D waypoint annotations. Experiments show that 3DG-VLN outperforms competitive UAV-VLN baselines, achieving a 13.82\% improvement in success rate. Real-world trials further demonstrate the potential of 3DG-VLN for practical see-and-reach navigation. The source code and benchmark are available at https://github.com/xuefanfu/3DG-VLN.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。