用单目深度估计提升农业机器人导航能力
MDE-AgriVLN: Agricultural Vision-and-Language Navigation with Monocular Depth Estimation
- 从单张图像生成深度特征,增强空间感知
- 在A2A基准上成功率提升至32%,路径误差降低至4.08米
- 适合农业机器人自主导航研究者参考
农业机器人在多种农事任务中发挥重要作用,但仍严重依赖人工操作或轨道系统进行移动。AgriVLN方法与A2A基准首次将视觉-语言导航(VLN)引入农业领域,使机器人能根据自然语言指令导航至目标位置。不同于人类双眼视觉,大多数农业机器人仅配备单摄像头实现单目视觉,导致空间感知能力受限。为此,本文提出基于单目深度估计的农业视觉-语言导航方法(MDE-AgriVLN),设计MDE模块从RGB图像生成深度特征,辅助多模态决策。在A2A基准上评估,该方法将成功率达由0.23提升至0.32,导航误差由4.43米降至4.08米,达到农业视觉-语言导航领域的最先进水平。
原文摘要 · Abstract (English)
Agricultural robots are serving as powerful assistants across a wide range of agricultural tasks, nevertheless, still heavily relying on manual operations or railway systems for movement. The AgriVLN method and the A2A benchmark pioneeringly extended Vision-and-Language Navigation (VLN) to the agricultural domain, enabling a robot to navigate to a target position following a natural language instruction. Unlike human binocular vision, most agricultural robots are only given a single camera for monocular vision, which results in limited spatial perception. To bridge this gap, we present the method of Agricultural Vision-and-Language Navigation with Monocular Depth Estimation (MDE-AgriVLN), in which we propose the MDE module generating depth features from RGB images, to assist the decision-maker on multimodal reasoning. When evaluated on the A2A benchmark, our MDE-AgriVLN method successfully increases Success Rate from 0.23 to 0.32 and decreases Navigation Error from 4.43m to 4.08m, demonstrating the state-of-the-art performance in the agricultural VLN domain. Code: https://github.com/AlexTraveling/MDE-AgriVLN.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。