arXiv:2506.17110cs.ROcs.CV2025-06中稿 · IROS 2025被引 5

用单张图片实现精准深度估计,无需额外数据即可完成机器人抓取

Monocular One-Shot Metric-Depth Alignment for RGB-Based Robot Grasping

  • 通过一次适应性校准,利用稀疏真值深度点对齐尺度旋转平移
  • 在真实场景中抓取成功率高,透明物体也能准确估计深度
  • 无需重训练模型,适合实际部署的机器人操作任务

精确的6D物体位姿估计是成功完成机器人抓取与非抓取操作的前提。目前主流方法依赖结构光、飞行时间或双目视觉等深度传感器,但存在成本高、噪声大、难以处理透明物体等问题。而现有单目深度估计模型仅提供仿射不变的深度(含未知尺度和偏移)。虽然部分模型在公开数据集上实现了零样本成功,但泛化能力差。本文提出一种新框架——单目一次式度量深度对齐(MOMA),通过一次适应性学习,从单张RGB图像恢复度量深度。MOMA在相机标定阶段利用稀疏真值深度点进行尺度-旋转-平移对齐,无需额外数据采集或模型重训练即可实现高精度深度估计。该方法支持对透明物体的微调,展现出强大泛化能力。真实世界实验在桌面双指抓取和吸盘式分拣任务中均取得高成功率,验证了其有效性。

原文摘要 · Abstract (English)

Accurate 6D object pose estimation is a prerequisite for successfully completing robotic prehensile and non-prehensile manipulation tasks. At present, 6D pose estimation for robotic manipulation generally relies on depth sensors based on, e.g., structured light, time-of-flight, and stereo-vision, which can be expensive, produce noisy output (as compared with RGB cameras), and fail to handle transparent objects. On the other hand, state-of-the-art monocular depth estimation models (MDEMs) provide only affine-invariant depths up to an unknown scale and shift. Metric MDEMs achieve some successful zero-shot results on public datasets, but fail to generalize. We propose a novel framework, Monocular One-shot Metric-depth Alignment (MOMA), to recover metric depth from a single RGB image, through a one-shot adaptation building on MDEM techniques. MOMA performs scale-rotation-shift alignments during camera calibration, guided by sparse ground-truth depth points, enabling accurate depth estimation without additional data collection or model retraining on the testing setup. MOMA supports fine-tuning the MDEM on transparent objects, demonstrating strong generalization capabilities. Real-world experiments on tabletop 2-finger grasping and suction-based bin-picking applications show MOMA achieves high success rates in diverse tasks, confirming its effectiveness.

机器人抓取单目深度姿态估计实时部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。