用点表示抓取姿态,提升机器人抓动态物体的稳定性与泛化能力
GAP-RL: Grasps As Points for RL Towards Dynamic Object Grasping
- 将6D抓取姿态转为高斯点,构建更抽象的抓取特征编码器
- 在仿真中对新物体和未知运动的抓取成功率超基线15%以上
- 适合需要实时抓取动态目标的工业机器人场景
复杂连续运动场景下对移动物体的动态抓取仍具挑战。强化学习(RL)因具备闭环特性被应用于多种机器人操作任务,但现有基于RL的方法未充分挖掘视觉表征的潜力。本文提出一种名为GAP-RL的新框架,用于高效可靠地抓取移动物体。通过快速区域级抓取检测器,将6D抓取位姿转换为高斯点,并提取抓取特征作为比原始物体点特征更高阶的抽象表示,构建抓取编码器。同时,设计了适用于真实部署的可抓取区域探索模块,搜索一致的可抓取区域,实现更平滑的抓取生成与稳定的策略执行。为公平评估性能,构建了一个包含多种复杂运动物体的仿真动态抓取基准。实验结果表明,本方法在新物体和未见动态运动上均表现出更强的泛化能力,优于其他基线方法。真实世界实验进一步验证了该框架的仿真到现实迁移能力。
原文摘要 · Abstract (English)
Dynamic grasping of moving objects in complex, continuous motion scenarios remains challenging. Reinforcement Learning (RL) has been applied in various robotic manipulation tasks, benefiting from its closed-loop property. However, existing RL-based methods do not fully explore the potential for enhancing visual representations. In this letter, we propose a novel framework called Grasps As Points for RL (GAP-RL) to effectively and reliably grasp moving objects. By implementing a fast region-based grasp detector, we build a Grasp Encoder by transforming 6D grasp poses into Gaussian points and extracting grasp features as a higher-level abstraction than the original object point features. Additionally, we develop a Graspable Region Explorer for real-world deployment, which searches for consistent graspable regions, enabling smoother grasp generation and stable policy execution. To assess the performance fairly, we construct a simulated dynamic grasping benchmark involving objects with various complex motions. Experiment results demonstrate that our method effectively generalizes to novel objects and unseen dynamic motions compared to other baselines. Real-world experiments further validate the framework's sim-to-real transferability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。