不依赖外设的机器人抓取校准,提升真实环境抓取成功率
Optimization of sim-to-real transfer in the humanoid robot NICO

- 基于视觉与手眼校准,构建无需外设的抓取流程
- 视觉反馈使全桌抓取成功率最高,校准区表现最优
- 适用于低成本人形机器人在复杂场景下的精准操作
机器人抓取需要视觉感知、目标定位、逆运动学和手部控制的精确协同。然而,仿真中规划的动作在物理机器人上执行时,由于仿真-现实差距会引发微小定位误差,导致抓取失败。此前我们提出一种低成本触觉校准方法,提升了人形机器人NICO的2D到达精度。本文将该方法从到达扩展至桌面物体抓取,新增基于YOLO的目标与手部检测、利用机器人内置低分辨率鱼眼相机的立体视觉定位,以及针对抓取任务的特定修正。这些组件共同构成一个无需RGB-D相机、动作捕捉或外部追踪系统的新型校准式抓取流程。同时,我们实现了视觉反馈模型,在抓取前对齐机械手与目标。结果表明,完全非线性校准模型在标定区域内表现最佳,而视觉反馈模型在整个桌面工作空间中实现了最高的整体抓取成功率。
原文摘要 · Abstract (English)
Robotic grasping requires accurate coordination between visual perception, object localization, inverse kinematics, and hand control. However, when movements planned in simulation are executed on a physical robot, the sim-to-real gap can cause small positioning errors that prevent successful grasping. In our previous work, we introduced a low-cost haptic calibration method that improved 2D reaching accuracy of the humanoid robot NICO. In this paper, we extend this approach from reaching to tabletop object grasping by adding YOLO-based object and hand detection, stereo vision-based localization using the robot's built-in low-resolution fisheye cameras, and task-specific corrections for grasp execution. Together, these components form a novel calibration-based grasping pipeline that does not require RGB-D cameras, motion capture, or external tracking systems. We also implemented a visual feedback model that aligns the robot hand with the detected object before grasping. Our results show that the fully nonlinear calibration model achieved the best performance inside the calibrated area, while the visual feedback model achieved the highest overall grasping success across the full tabletop workspace.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。