用视觉和身体感知推断接触,让机器人无需触觉传感器也能灵巧操作物体。
NoContactNoWorries: Estimating Contact through Vision and Proprioception for In-Hand Dexterous Manipulation

- 融合RGB-D图像与机器人本体感觉,用Transformer模型预测接触状态。
- 在仿真和真实机器人上均实现对多种物体的接触准确推断。
- 适合缺乏触觉硬件但需灵巧操作的机器人系统使用。
精确感知物理接触是灵巧操作的基础。尽管机器人常依赖专用触觉传感器,人类却能通过视觉与身体姿态、运动的内在感知来推断接触。受此启发,我们探索是否可让机器人仅通过视觉与本体感觉学习推断接触,这为二值接触估计提供了一种可扩展的替代方案,避免了触觉硬件在成本、脆弱性和集成上的实际挑战。我们提出NoContactNoWorries,一种基于Transformer的多模态框架,融合RGB-D视觉与机器人本体感觉,生成用于手-物交互的伪触觉信号。我们在多个物体上训练单一接触预测模型,结果表明该推断信号可支持下游强化学习代理完成物体在手中重定向任务,并泛化至新物体。仿真与真实机器人实验验证了该方法的可行性,证明了从视觉与本体感觉中推断接触的潜力。
原文摘要 · Abstract (English)
Perceiving physical contact is fundamental to dexterous manipulation. While robots often rely on dedicated hardware tactile sensors, humans exhibit a remarkable ability to infer contact by integrating visual information with an innate sense of their body's pose and movement. Inspired by this embodied perceptual skill, we investigate whether a robot can learn to infer contact from vision, an approach that also offers a scalable alternative to tactile hardware specifically for binary contact estimation, which faces practical challenges in cost, fragility, and integration. We present NoContactNoWorries, a transformer-based multimodal framework that fuses RGB-D vision with the robot's proprioception to infer binary contact states as a pseudo-tactile signal for hand-object interactions. We validate by training a single contact prediction model on multiple objects and show that the inferred contact signal supports downstream reinforcement learning agents for in-hand object reorientation, generalizing to novel objects. Experiments in both simulation and on a real-world robot validate our approach, highlighting the feasibility of inferring contact from vision and proprioception. Project Page: https://soham2560.github.io/no-contact-no-worries/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。