基于根部几何的抓取框架,实现果蔬稳定定向抓取与放置
ROG-Grasp: Root-Oriented Geometry for Robotic Grasping and Placement

- 通过根部表面几何估计物体朝向,结合RGB-D感知生成稳定抓取姿态
- 在番茄和洋葱上实现超过90%的成功率,且在杂乱场景中执行时间稳定
- 适合需要精确朝向控制的农业采摘场景,优于视觉-语言-动作模型
面向收获后农产品处理中的定向操作需求,本文提出ROG-Grasp——一种基于几何的机器人抓取与放置框架。该方法利用RGB-D感知从根部表面几何推断作物朝向,采用基于YOLO的根部检测与点云平面拟合计算根法向量,从而生成稳定的抓取位姿并实现朝向约束的笛卡尔运动规划。在番茄和洋葱上的实验表明,该方法在孤立与杂乱场景下均实现高成功率(>90%)与稳定的执行时间。相比视觉-语言-动作(VLA)策略,本方法在抓取完成率与执行速度上表现更优,验证了几何驱动感知在实际定向操作任务中的有效性。视频演示见https://youtu.be/Ir2UtGODdMo。
原文摘要 · Abstract (English)
Orientation-aware manipulation is essential in post-harvest agricultural processing, where produce must be grasped and placed in consistent configurations. This paper presents ROG-Grasp, a geometry-based robotic grasping and placement framework that estimates the produce orientation from root surface geometry using RGB-D perception. A YOLO-based root detector and point cloud plane fitting are used to infer the root normal, enabling stable grasp pose generation and orientation-constrained Cartesian motion planning. Experiments on tomatoes and onions demonstrate high success rates and stable execution time in both isolated and cluttered scenarios. Compared with vision-language-action (VLA) policies, the proposed method achieves more reliable and accurate grasp completion with faster execution. These results highlight the effectiveness of geometry-driven perception for practical orientation-controlled manipulation tasks. A video of our paper is available online https://youtu.be/Ir2UtGODdMo.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。