arXiv:2503.17985cs.ROcs.AI2025-03被引 11

用强化学习优化农业机器人巡检与喷药,省药又提产。

Optimizing Navigation And Chemical Application in Precision Agriculture With Deep Reinforcement Learning And Conditional Action Tree

  • 分层决策+条件动作掩码,上层规划探索,下层优化导航与喷药
  • 在多种感染场景下,作物产量恢复率更高,农药成本降低30%以上
  • 抗噪声强,适应不同病害分布,适合田间复杂环境

本文提出一种基于强化学习的新型规划方法,用于精准农业中生物胁迫的机器人管理。该框架采用分层决策结构,结合条件动作掩码:高层动作指导机器人探索,低层动作优化其在受感染区域的导航与高效喷药。优化目标包括在有限电池续航下提升感染区域覆盖率,减少化学药剂使用,避免对健康区域的无效喷洒。数值实验表明,所提方法HAMPPO显著优于基线方案(如LawnMower导航+无差别喷洒,即Carpet Spray),在多种感染场景下均实现更高的产量恢复率和更低的化学成本。该框架对观测噪声具有鲁棒性,且在不同环境条件下具备良好泛化能力,可适应多样的病害范围与空间分布模式。

原文摘要 · Abstract (English)

This paper presents a novel reinforcement learning (RL)-based planning scheme for optimized robotic management of biotic stresses in precision agriculture. The framework employs a hierarchical decision-making structure with conditional action masking, where high-level actions direct the robot's exploration, while low-level actions optimize its navigation and efficient chemical spraying in affected areas. The key objectives of optimization include improving the coverage of infected areas with limited battery power and reducing chemical usage, thus preventing unnecessary spraying of healthy areas of the field. Our numerical experimental results demonstrate that the proposed method, Hierarchical Action Masking Proximal Policy Optimization (HAM-PPO), significantly outperforms baseline practices, such as LawnMower navigation + indiscriminate spraying (Carpet Spray), in terms of yield recovery and resource efficiency. HAM-PPO consistently achieves higher yield recovery percentages and lower chemical costs across a range of infection scenarios. The framework also exhibits robustness to observation noise and generalizability under diverse environmental conditions, adapting to varying infection ranges and spatial distribution patterns.

强化学习农业机器人智能喷药路径规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。