arXiv:2506.12239cs.ROcs.CV2025-06中稿 · RSS 2025 | Project…被引 13

融合视觉与触觉,精准估计物体在手姿态与外部接触点。

ViTaSCOPE: Visuo-tactile Implicit Representation for In-hand Pose and Extrinsic Contact Estimation

  • 用隐式神经场融合视觉与高分辨率触觉信号
  • 在仿真中训练,实现零样本迁移至真实世界
  • 适合需要精细操作的机器人抓取任务

灵巧操作要求精确估计物体在手姿态和外部接触位置,但因观测不完整且噪声大,难度很高。我们提出ViTaSCOPE:一种以物体为中心的视觉-触觉联合隐式表征方法,将物体表示为有符号距离场,将分布式触觉反馈建模为神经剪切场,从而精确定位物体并将其外部接触注册到3D几何上形成接触场。通过利用仿真进行可扩展训练,并有效弥合仿真到现实的差距,该方法实现了对互补视觉-触觉线索的无缝推理。我们在大量模拟和真实世界实验中验证了其在灵巧操作场景中的能力。

原文摘要 · Abstract (English)

Mastering dexterous, contact-rich object manipulation demands precise estimation of both in-hand object poses and external contact locations$\unicode{x2013}$tasks particularly challenging due to partial and noisy observations. We present ViTaSCOPE: Visuo-Tactile Simultaneous Contact and Object Pose Estimation, an object-centric neural implicit representation that fuses vision and high-resolution tactile feedback. By representing objects as signed distance fields and distributed tactile feedback as neural shear fields, ViTaSCOPE accurately localizes objects and registers extrinsic contacts onto their 3D geometry as contact fields. Our method enables seamless reasoning over complementary visuo-tactile cues by leveraging simulation for scalable training and zero-shot transfers to the real-world by bridging the sim-to-real gap. We evaluate our method through comprehensive simulated and real-world experiments, demonstrating its capabilities in dexterous manipulation scenarios.

机器人操作触觉感知隐式表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。