让农业机器人听懂自然语言指令,在农田中自主导航。
AgriVLN: Vision-and-Language Navigation for Agricultural Robots
- 用视觉语言模型理解指令和农田环境,生成控制动作
- 在6类农田场景中成功率达47%,较基线提升14个百分点
- 专为农业设计的基准与导航系统,适合智能农机研发者
农业机器人虽具潜力,但普遍依赖人工操作或固定轨道,移动性差、适应性弱。视觉-语言导航(VLN)使机器人可依自然语言指令自主导航,在多个领域表现优异,但现有方法均未针对农业场景设计。为此,我们提出农业到农业(A2A)基准,包含6种真实农田场景、1560个导航任务,所有图像由高0.38米的四足机器人前视摄像头拍摄,贴近实际部署条件。同时提出面向农业机器人的视觉-语言导航(AgriVLN)基线,基于视觉语言模型(VLM),通过精心设计的提示模板理解指令与环境,生成低层控制动作。在A2A上评估显示,该方法对短指令表现良好,但长指令失败率高,因难以追踪当前执行指令部分。为此,我们引入子任务列表(STL)分解模块,将指令拆解为可追踪子任务,使成功率从0.33提升至0.47。此外,与多种现有VLN方法对比,证明其在农业领域达到先进水平。
原文摘要 · Abstract (English)
Agricultural robots have emerged as powerful members in agricultural tasks, nevertheless, still heavily rely on manual operation or untransportable railway for movement, resulting in limited mobility and poor adaptability. Vision-and-Language Navigation (VLN) enables robots to navigate to the target destinations following natural language instructions, demonstrating strong performance on several domains. However, none of the existing benchmarks or methods is specifically designed for agricultural scenes. To bridge this gap, we propose Agriculture to Agriculture (A2A) benchmark, containing 1,560 episodes across six diverse agricultural scenes, in which all realistic RGB videos are captured by front-facing camera on a quadruped robot at a height of 0.38 meters, aligning with the practical deployment conditions. Meanwhile, we propose Vision-and-Language Navigation for Agricultural Robots (AgriVLN) baseline based on Vision-Language Model (VLM) prompted with carefully crafted templates, which can understand both given instructions and agricultural environments to generate appropriate low-level actions for robot control. When evaluated on A2A, AgriVLN performs well on short instructions but struggles with long instructions, because it often fails to track which part of the instruction is currently being executed. To address this, we further propose Subtask List (STL) instruction decomposition module and integrate it into AgriVLN, improving Success Rate (SR) from 0.33 to 0.47. We additionally compare AgriVLN with several existing VLN methods, demonstrating the state-of-the-art performance in the agricultural domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。