用分层强化学习让机器人在草莓丛中精准分离并采摘成熟果实。
Vision-Based Obstacle Separation for Strawberry Harvesting in Clusters Using Hierarchical Reinforcement Learning

- 分两阶段处理:先视觉引导分离障碍物,再精准抓取目标。
- 仿真成功率96.7%,实机成功率达71.7%~88.3%。
- 适合复杂遮挡环境下的农业机器人采摘任务。
在密集草莓簇中进行选择性采摘面临挑战,因成熟果实常被周围未成熟果实遮挡,直接抓取不可靠。本文提出一种分层强化学习框架VGPA,融合视觉引导决策机制与渐进式自适应探索策略(PAES),实现基于视觉的障碍物分离与采摘。任务分为两个连续阶段:障碍物分离和目标抓取。高层采用视觉引导机制,提升选项选择能力并加速策略收敛;底层使用PAES,提高连续控制学习中的探索效率与训练稳定性。仿真实验显示,所学策略成功率高达96.7%。此外,在自研并联机器人上的模拟到现实迁移实验表明,该方法成功率在71.7%至88.3%之间,优于直接采摘,仅多耗时1.22~1.4秒。结果验证了该方法在复杂簇状环境下的有效性、泛化能力与实际应用潜力。
原文摘要 · Abstract (English)
Selective harvesting in clustered strawberry environments is challenging because ripe fruits are often occluded by surrounding unripe fruits, making direct grasping unreliable. To address this problem, this paper proposes a hierarchical reinforcement learning framework, termed VGPA, which integrates a vision-guided decision mechanism and a Progressive Adaptive Exploration Strategy (PAES) for vision-based obstacle separation and harvesting. The task was decomposed into two sequential stages: obstacle separation and target grasping. At the high level, the vision-guided mechanism improved option selection and accelerated policy convergence. At the low level, PAES improved exploration efficiency and training stability during continuous control learning. In simulation experiments, the learned policy achieved a success rate of 96.7%. In addition, sim-to-real transfer experiments on a self-developed parallel robot showed that the proposed method achieved success rates ranging from 71.7% to 88.3%, outperforming direct picking while requiring only 1.22~s more average harvesting time. These results verified the effectiveness, generalization ability, and practical potential of the proposed method for robotic harvesting in complex clustered environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。