单摄像头实现视觉触觉同步感知,突破传统传感器遮挡限制。
MuxGel: Simultaneous Dual-Modal Visuo-Tactile Sensing via Spatially Multiplexing and Deep Reconstruction

- 通过棋盘纹涂层实现视觉与触觉区域空间复用
- 重建网络在未见物体上实现高精度双模态信号还原
- 适用于机械手抓取任务,适合需要多模态反馈的场景
高保真视觉-触觉感知对精准机器人操作至关重要,但多数基于视觉的触觉传感器依赖不透明涂层,会遮挡外部视觉观察。本文提出MuxGel,一种通过单个相机同时捕捉外部视觉信息和接触引发的触觉信号的空间复用传感器。采用棋盘状涂层设计,将敏感触觉区域与透明窗口交错排列,既保留标准外形,又支持即插即用式集成至GelSight型传感器中。为从复用输入中恢复密集的视觉与触觉信号,我们构建了一个基于U-Net的重建框架,采用模拟到真实训练流程进行优化。在未见过的物体上的实验验证了该方法的泛化能力与精度。进一步在抓取任务中展示:视觉反馈用于对齐,触觉反馈用于接触交互。结果表明,MuxGel可在GelSight型结构中实现单相机双模态感知,提供局部视觉反馈与重建触觉反馈,具备向其他光学触觉传感器扩展的潜力。
原文摘要 · Abstract (English)
High-fidelity visuo-tactile sensing is important for precise robotic manipulation, yet most vision-based tactile sensors rely on opaque coatings that enable tactile sensing but block direct visual observation. We propose MuxGel, a spatially multiplexed sensor that captures both external visual information and contact-induced tactile signals through a single camera. By using a checkerboard coating pattern, MuxGel interleaves tactile-sensitive regions with transparent windows for external vision. This design maintains standard form factors, allowing for plug-and-play integration into GelSight-style sensors by simply replacing the gel pad. To recover dense visual and tactile signals from the multiplexed inputs, we develop a U-Net-based reconstruction framework trained with a sim-to-real pipeline. Experiments on unseen objects demonstrate the framework's generalization and accuracy. We further demonstrate MuxGel in grasping tasks, where visual feedback supports alignment and tactile feedback supports contact interaction. Results show that MuxGel enables single-camera dual-modal sensing within a GelSight-style implementation, providing local visual feedback and reconstructed tactile feedback with potential extension to other optical tactile sensors. Project webpage: https://zhixianhu.github.io/muxgel/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。