用模仿学习提升无人机农田巡检效率,降低能耗并提高识别率。
Trajectory Planning for UAV-Based Smart Farming Using Imitation-Based Triple Deep Q-Learning
- 引入模仿学习的三重深度Q网络,减少探索成本。
- 实测识别率提升4.43%,数据采集率提升6.94%。
- 适合农业无人机路径规划与强化学习初学者参考。
无人飞行器(UAV)已成为智慧农业的有力辅助平台,可同步完成杂草检测、识别及无线传感器数据采集。然而,由于环境高度不确定、观测不完整以及飞行器电池容量有限,基于UAV的路径规划面临挑战。为此,本文将路径规划问题建模为马尔可夫决策过程(MDP),并采用多智能体强化学习(MARL)求解。进一步提出一种新型基于模仿学习的三重深度Q网络(ITDQN)算法,通过精英模仿机制降低探索开销,并在双深度Q网络(DDQN)基础上引入中介Q网络,加速训练并提升性能稳定性。在仿真与真实环境中的实验结果表明,所提方法有效。ITDQN相比DDQN在杂草识别率上提升4.43%,数据采集率提升6.94%。
原文摘要 · Abstract (English)
Unmanned aerial vehicles (UAVs) have emerged as a promising auxiliary platform for smart agriculture, capable of simultaneously performing weed detection, recognition, and data collection from wireless sensors. However, trajectory planning for UAV-based smart agriculture is challenging due to the high uncertainty of the environment, partial observations, and limited battery capacity of UAVs. To address these issues, we formulate the trajectory planning problem as a Markov decision process (MDP) and leverage multi-agent reinforcement learning (MARL) to solve it. Furthermore, we propose a novel imitation-based triple deep Q-network (ITDQN) algorithm, which employs an elite imitation mechanism to reduce exploration costs and utilizes a mediator Q-network over a double deep Q-network (DDQN) to accelerate and stabilize training and improve performance. Experimental results in both simulated and real-world environments demonstrate the effectiveness of our solution. Moreover, our proposed ITDQN outperforms DDQN by 4.43\% in weed recognition rate and 6.94\% in data collection rate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。