用仿真训练的强化学习+鲁棒视觉,实现高成功率草莓采摘。
Robotic Strawberry Harvesting with Robust Vision and Deep Reinforcement Learning based Sim-to-Real Control

- 结合改进的视觉模型与仿真训练的强化学习策略
- 实测采摘成功率96.6%,整体成功率达84.3%
- 适合农业机器人研发人员参考落地
本研究提出一种闭环式机器人草莓采摘系统,融合鲁棒视觉模块、基于模拟训练的深度强化学习(DRL)控制以及基于ROS的真实机器人执行。感知方面,提出HRAttnEdge-YOLO26-seg,通过高分辨率P2分支、分割路径注意力和边缘监督原型学习,提升复杂场景下的实例分割性能。控制方面,在Isaac Lab中训练目标条件化的近端策略优化(PPO)策略,生成UR10e机械臂平滑关节位置指令,并部署于真实机器人完成目标果实抓取与采摘。该仿真驱动方法降低硬件依赖,减少开发成本,支持无大量物理试验的可扩展策略训练。所提视觉模型在自建及公开数据集上均表现最优,分割性能提升10至14%。室内测试中,PPO控制器运动更稳定平滑,优于基于逆运动学(IK)的MoveIt基线。温室实测共采摘281颗草莓,达到96.6%的到达成功率、91.3%的抓取-拉摘成功率和84.3%的整体采摘成功率。结果表明,任务专用感知结合仿真训练的PPO可作为传统规划依赖抓取的有效替代方案,实现复杂农作环境中的可靠闭环采摘。
原文摘要 · Abstract (English)
This study presents a closed-loop robotic strawberry harvesting system that combines a robust vision module, simulation-trained deep reinforcement learning (DRL) control, and ROS-based realrobot execution. For perception, we propose HRAttnEdge-YOLO26-seg, a modified YOLO26-seg architecture that incorporates a high-resolution P2 branch, segmentation-path attention, and edgesupervised prototype learning to improve instance segmentation in cluttered scenes. For control, we train a target-conditioned Proximal Policy Optimization (PPO) policy in Isaac Lab to produce smooth joint-position commands for a UR10e manipulator and deploy it on a UR10e robot for targetfruit reaching and harvesting. This simulation-based approach reduces hardware dependency, lowers development cost, and allows scalable policy training without exhaustive physical trials before real deployment. The proposed vision model demonstrated the highest overall performance among the evaluated methods. On both self-collected and public datasets, the model showed a 10 to 14% improvement in segmentation performance. In controlled in-house tests, the PPO controller produced stable and dynamically smoother motion than a inverse kinematics (IK)-based MoveIt baseline. In greenhouse trials, the proposed integrated system harvested 281 strawberries, achieving 96.6% reaching success, 91.3% grasp-and-pull success, and 84.3% overall harvesting success. These results illustrate that task-specific perception combined with simulation-trained PPO can serve as a practical and resource-efficient alternative to conventional planner-dependent reaching in manipulation, enabling reliable closed-loop robotic harvesting in complex agricultural environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。