用视觉语言模型指导农业机器人规划,实现精准作物监测
Visual-Language-Guided Task Planning for Horticultural Robots
- 通过视觉语言模型主动查询多源数据,动态规划机器人任务
- 短时任务成功率87%,长时多目标任务在无噪声下仍超76%
- 揭示了当前模型对语义地图噪声敏感的关键缺陷
作物监测对精准农业至关重要,但现有系统缺乏高层推理能力。本文提出一种新颖的模块化框架,利用视觉语言模型(VLM)通过主动查询异构数据源(包括增强的RGB图像和2D语义占用图)并结合机器人动作原语,指导机器人任务规划。我们构建了一个涵盖单作与混作环境的短时与长时作物监测任务综合基准。结果显示,零样本VLM在短时任务中表现稳健(成功率87%,接近人类专家水平),但在复杂长时多目标任务中成功率降至10%以下;然而在无噪声条件下,任务完成率仍高于76%。关键发现是:系统在依赖噪声语义地图时性能显著下降,暴露出当前VLM在持续机器人操作中的上下文定位局限性。本工作提供可部署框架,并深入揭示了VLM在复杂农业机器人应用中的能力与瓶颈。
原文摘要 · Abstract (English)
Crop monitoring is essential for precision agriculture, but current systems lack high-level reasoning. We introduce a novel, modular framework that uses a Vision Language Model (VLM) to guide robotic task planning by actively querying heterogeneous data sources, including enriched RGB camera feeds and 2D semantic occupancy maps, interleaved with robotic action primitives. We contribute a comprehensive benchmark for short- and long-horizon crop monitoring tasks in monoculture and polyculture environments. Our results show that while zero-shot VLMs perform robustly for short-horizon tasks (achieving 87% success, comparable to human experts), success drops significantly to under 10% for complex long-horizon, multi-target tasks. Despite this decline, task completion rates remain above 76% under noiseless conditions. Critically, the system degrades when relying on noisy semantic maps, demonstrating a key limitation in current VLM context grounding for sustained robotic operations. This work offers a deployable framework and critical insights into VLM capabilities and shortcomings for complex agricultural robotics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。