arXiv:2601.01675cs.RO2026-01中稿 · ICRA被引 70

融合视觉与触觉数据,提升机器人抓握物体的6维姿态估计精度。

VisuoTactile 6D Pose Estimation of an In-Hand Object using Vision and Tactile Sensor Data

  • 用点云表示触觉传感器接触的物体表面,实现像素级密集融合。
  • 触觉+视觉联合估计使6D姿态误差降低,合成数据训练可迁移到真实机器人。
  • 适合研究多模态感知、机器人灵巧操作的开发者参考。

掌握物体的6维姿态有助于实现抓握中的精细操作。由于机械手遮挡严重,仅依赖视觉的方法效果受限。许多机器人在指尖配备触觉传感器,可补充视觉信息。本文提出一种结合视觉与触觉数据的在手物体6维姿态估计方法。针对触觉数据缺乏标准表示及传感器融合难题,采用点云描述触觉接触区域,并设计基于像素级密集融合的网络架构。同时扩展NVIDIA深度学习数据集生成器,生成逼真的合成视觉数据与对应触觉点云。实验表明,引入触觉数据显著提升了6维姿态估计性能,且模型从合成数据训练成功迁移到真实物理机器人上。

原文摘要 · Abstract (English)

Knowledge of the 6D pose of an object can benefit in-hand object manipulation. In-hand 6D object pose estimation is challenging because of heavy occlusion produced by the robot's grippers, which can have an adverse effect on methods that rely on vision data only. Many robots are equipped with tactile sensors at their fingertips that could be used to complement vision data. In this paper, we present a method that uses both tactile and vision data to estimate the pose of an object grasped in a robot's hand. To address challenges like lack of standard representation for tactile data and sensor fusion, we propose the use of point clouds to represent object surfaces in contact with the tactile sensor and present a network architecture based on pixel-wise dense fusion. We also extend NVIDIA's Deep Learning Dataset Synthesizer to produce synthetic photo-realistic vision data and corresponding tactile point clouds. Results suggest that using tactile data in addition to vision data improves the 6D pose estimate, and our network generalizes successfully from synthetic training to real physical robots.

姿态估计触觉融合多模态感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。