融合视觉与触觉的Transformer模型,解决精密装配中的微米级对准难题。
ReTac-ACT: A State-Gated Vision-Tactile Fusion Transformer for Precision Assembly
- 双向交叉注意力增强视觉与触觉特征,实现互惠互补。
- 在视觉遮挡时动态提升触觉依赖度,0.1mm间隙下仍达80%成功率。
- 专为工业级高精度装配设计,适合机器人抓取与自动化制造场景。
精密装配需要在接触密集的‘最后一毫米’区域进行亚毫米级调整,而传统视觉反馈因末端执行器和工件遮挡失效。我们提出ReTac-ACT(重建增强的触觉-视觉融合变压器),一种基于模仿学习的视觉-触觉策略,通过三种协同机制解决该问题:(i) 双向交叉注意力实现融合前的视觉与触觉特征互增强;(ii) 前馈感知条件门控网络,在视觉遮挡时动态提升触觉依赖性;(iii) 触觉重建目标,强制学习与操作相关的接触信息而非通用纹理。在标准NIST Assembly Task Board M1基准上,ReTac-ACT达到90%的插销成功率,显著优于纯视觉及通用基线方法,并在工业级0.1mm间隙下保持80%成功率。消融实验验证每个模块均不可或缺。代码库与包含多种间隙水平的视觉-触觉演示数据集将公开,支持可复现研究。
原文摘要 · Abstract (English)
Precision assembly requires sub-millimeter corrections in contact-rich "last-millimeter" regions where visual feedback fails due to occlusion from the end-effector and workpiece. We present ReTac-ACT (Reconstruction-enhanced Tactile ACT), a vision-tactile imitation learning policy that addresses this challenge through three synergistic mechanisms: (i) bidirectional cross-attention enabling reciprocal visuo-tactile feature enhancement before fusion, (ii) a proprioception-conditioned gating network that dynamically elevates tactile reliance when visual occlusion occurs, and (iii) a tactile reconstruction objective enforcing learning of manipulation-relevant contact information rather than generic visual textures. Evaluated on the standardized NIST Assembly Task Board M1 benchmark, ReTac-ACT achieves 90% peg-in-hole success, substantially outperforming vision-only and generalist baseline methods, and maintains 80% success at industrial-grade 0.1mm clearance. Ablation studies validate that each architectural component is indispensable. The ReTac-ACT codebase and a vision-tactile demonstration dataset covering various clearance levels with both visual and tactile features will be released to support reproducible research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。