arXiv:2606.08765cs.ROcs.CV2026-06被引 1

将触觉信息投影到图像空间,提升机器人在视觉遮挡下的灵巧操作能力。

RGB-S: Image-Aligned Tactile Saliency for Robust Dexterous Manipulation

论文配图:RGB-S: Image-Aligned Tactile Saliency for Robust Dexterous Manipulation
图 1 · 摘自论文原文
  • 通过机械臂运动学和相机标定,将触觉传感器位置映射到图像平面。
  • 在真实场景中,遮挡条件下操作成功率提升26.7个百分点。
  • 适合需要强鲁棒性视觉-触觉融合的机器人操控任务。

有效的视觉-触觉融合对机器人灵巧操作至关重要,尤其在视觉观测不可靠或被遮挡时。然而,如何将稀疏、异构的触觉测量与密集的视觉表征可靠对齐仍是一个基本挑战。现有方法通常依赖策略从有限演示中隐式学习跨模态对应关系,未利用几何先验,导致数据效率低且在视觉退化时泛化能力差。为此,我们提出一种显式将物理接触定位在图像域中的框架。利用机器人正向运动学和相机标定,我们将触觉传感器位置直接投影至RGB图像平面,并生成受力调制的高斯显著性图以建模由运动学和标定误差引起的时空不确定性。通过零初始化的条件架构整合这些二维空间锚点,我们的方法在不破坏预训练视觉表示的前提下,将物理接触先验注入标准视觉主干网络。我们在模拟和真实世界中六个灵巧操作任务上进行了评估,极端视觉遮挡条件下,实机实验表明,显式地在图像域中进行RGB-S对齐,使操作成功率相比最强的隐式视觉-触觉基线提升了26.7个百分点,表明其具备更强的空间推理能力和对遮挡的鲁棒性。项目页面:touch-as-saliency.github.io

原文摘要 · Abstract (English)

Effective visuo-tactile integration is critical for robotic dexterous manipulation, especially when visual observations are unreliable or occluded. However, robustly aligning sparse, heterogeneous tactile measurements with dense visual representations remains a fundamental challenge. Most existing approaches require policies to learn cross-modal correspondences implicitly from limited demonstrations, without leveraging geometric priors. As a result, they are often data-inefficient and generalize poorly when visual observations are degraded. To address this limitation, we propose a framework that explicitly grounds physical contacts in the image domain. Using robot forward kinematics and camera calibration, we project tactile sensor locations directly onto the RGB image plane. We then render force-modulated Gaussian saliency maps to model spatial uncertainty arising from kinematic and calibration errors. By integrating these 2D spatial anchors through a zero-initialized conditioning architecture, our method injects physical contact priors into standard visual backbones while preserving pre-trained visual representations. We evaluate our method on six dexterous manipulation tasks in both simulation and the real world under severe visual occlusions. Real-world experiments show that explicit RGB-S grounding in the image domain improves real-world occluded manipulation success rates by $26.7$ percentage points over the strongest implicit visuo-tactile baseline, suggesting its improved spatial reasoning and robustness to occlusion. Project page: touch-as-saliency.github.io

触觉融合灵巧操作视觉-触觉鲁棒控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。