用视觉生成触觉信号,让机器人无需真实触觉传感器就能精准推物。
ViTacGen: Robotic Pushing with Vision-to-Touch Generation
- 通过视觉序列生成接触深度图,模拟触觉反馈。
- 在仿真与真实场景中实现最高86%的推物成功率。
- 适合无触觉传感器的机器人系统,支持零样本部署。
机器人推动物体是基础操作任务,需触觉反馈捕捉末端执行器与物体间的细微作用力和动态。然而真实触觉传感器常受限于高成本、易损及校准困难等问题,而纯视觉策略性能不佳。受人类通过视觉推断触觉能力启发,本文提出ViTacGen框架,基于强化学习实现视觉到触觉的生成,无需高分辨率真实触觉传感器,可直接在仅依赖视觉的机器人系统上实现零样本部署。ViTacGen包含一个编码器-解码器结构的视觉-触觉生成网络,从视觉图像序列生成接触深度图(标准化触觉表示),再通过融合视觉与生成触觉信息的强化学习策略,结合对比学习提升感知能力。实验在仿真与真实世界均验证了方法有效性,成功率达86%。
原文摘要 · Abstract (English)
Robotic pushing is a fundamental manipulation task that requires tactile feedback to capture subtle contact forces and dynamics between the end-effector and the object. However, real tactile sensors often face hardware limitations such as high costs and fragility, and deployment challenges involving calibration and variations between different sensors, while vision-only policies struggle with satisfactory performance. Inspired by humans' ability to infer tactile states from vision, we propose ViTacGen, a novel robot manipulation framework designed for visual robotic pushing with vision-to-touch generation in reinforcement learning to eliminate the reliance on high-resolution real tactile sensors, enabling effective zero-shot deployment on visual-only robotic systems. Specifically, ViTacGen consists of an encoder-decoder vision-to-touch generation network that generates contact depth images, a standardized tactile representation, directly from visual image sequence, followed by a reinforcement learning policy that fuses visual-tactile data with contrastive learning based on visual and generated tactile observations. We validate the effectiveness of our approach in both simulation and real world experiments, demonstrating its superior performance and achieving a success rate of up to 86\%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。