arXiv:2505.03702cs.ROcs.CV2025-05被引 3

用自监督学习融合几何与神经网络,提升机器人抓取叶片成功率。

Self-Supervised Learning for Robotic Leaf Manipulation: A Hybrid Geometric-Neural Approach

  • 结合YOLOv8和RAFT-Stereo生成叶形特征,通过置信度动态融合几何与神经方法。
  • 在温室中实现84.7%抓取成功率,优于纯几何(75.3%)和纯神经方法(60.2%)。
  • 适合农业机器人、自动化种植系统研发人员参考,兼具精度与泛化能力。

在农业场景中自动化叶片操作面临植物形态多样性和叶片可变形的挑战。本文提出一种混合几何-神经的自监督学习方法,实现自主叶片抓取。该方法利用YOLOv8进行实例分割,RAFF-Stereo进行三维深度估计,构建丰富的叶形表征,并输入几何特征评分流程与神经精修模块GraspPointCNN。关键创新在于置信度加权融合机制,根据预测确定性动态调整各方法贡献。自监督框架以几何流程为专家教师,自动生成训练数据。实验表明,该方法在受控环境下成功率达88.0%,在真实温室条件下达84.7%,显著优于纯几何方法(75.3%)和纯神经方法(60.2%)。本工作为农业机器人开辟新范式,实现领域知识与机器学习能力的无缝融合,为全自动作物监测系统奠定基础。

原文摘要 · Abstract (English)

Automating leaf manipulation in agricultural settings faces significant challenges, including the variability of plant morphologies and deformable leaves. We propose a novel hybrid geometric-neural approach for autonomous leaf grasping that combines traditional computer vision with neural networks through self-supervised learning. Our method integrates YOLOv8 for instance segmentation and RAFT-Stereo for 3D depth estimation to build rich leaf representations, which feed into both a geometric feature scoring pipeline and a neural refinement module (GraspPointCNN). The key innovation is our confidence-weighted fusion mechanism that dynamically balances the contribution of each approach based on prediction certainty. Our self-supervised framework uses the geometric pipeline as an expert teacher to automatically generate training data. Experiments demonstrate that our approach achieves an 88.0% success rate in controlled environments and 84.7% in real greenhouse conditions, significantly outperforming both purely geometric (75.3%) and neural (60.2%) methods. This work establishes a new paradigm for agricultural robotics where domain expertise is seamlessly integrated with machine learning capabilities, providing a foundation for fully automated crop monitoring systems.

农业机器人自监督学习叶片抓取多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。