用视觉点云和触觉数据定位接触点,提升机器人感知精度。
VTLoc: Learning-based Tactile Contact Localization in Visual Point Clouds

- 融合视觉与触觉特征重建伪点云,实现空间对齐。
- 通过迭代更新,将接触点定位误差显著降低。
- 适用于真实物体的单点触觉定位,适合机器人抓取场景。
视觉与触觉是机器人感知与操作中互补的重要模态:视觉提供全局物体上下文,触觉则在接触点提供精确局部信息。将两者结合进行接触点定位——即预测触碰在物体表面的位置——面临巨大挑战,主要源于触觉数据与视觉几何之间的精确空间对齐需求。为此,我们提出VTLoc,一种基于视觉-触觉的新型框架,利用3D点云作为视觉输入,从触觉读数中定位接触点。VTLoc引入两个关键组件:几何多模态对齐模块,通过融合视觉-触觉特征重建伪点云,并将其与原始视觉点云对齐,以强制跨模态的空间一致性;以及迭代定位更新器,利用融合后的特征逐步精炼接触位置预测。在包含100个真实物体的新基准上评估,VTLoc有效减少了单点触觉定位中的局部-全局对应歧义,显著提升了定位准确率。
原文摘要 · Abstract (English)
Vision and touch are complementary modalities essential for robotic perception and manipulation. While vision provides global object context, touch offers precise local information at contact points. Integrating these modalities for contact localization, i.e., predicting the location of touch on an object's surface, poses significant challenges due to the need for accurate spatial alignment between tactile data and visual geometry. To address this challenge, we propose VTLoc, a novel visual-tactile framework that localizes contact points from tactile readings using a 3D point cloud as visual input. VTLoc introduces two key components: a geometric multi-modal alignment module, which reconstructs a pseudo-point cloud from fused visual-tactile features and aligns it with the visual point cloud to enforce spatial consistencies across modalities; and an iterative localizing updater, which iteratively refines the predicted contact location using fused visual-tactile features. Evaluated on a new benchmark of 100 real-world objects, VTLoc improves single-touch contact localization by reducing local-to-global correspondence ambiguity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。