用物理约束提升触觉与视觉融合,机器人插接成功率超96%
Symmetry-Aware Fusion of Vision and Tactile Sensing via Bilateral Force Priors for Robotic Manipulation
- 设计跨模态变压器,通过自注意力与交叉注意力融合视觉和触觉信号
- 引入双侧受力平衡正则化,使触觉特征更稳定,插接成功率达96.59%
- 适合做精密操作的机器人系统,尤其关注触觉反馈的多模态研究者
机器人插接任务需要精确且接触丰富的交互,仅靠视觉无法解决。尽管触觉反馈直观有用,但现有方法的简单视觉-触觉融合常无法带来稳定提升。本文提出一种跨模态变压器(CMT),通过结构化自注意力和交叉注意力,将腕部相机观测与触觉信号融合。为稳定触觉嵌入,进一步引入物理启发的正则化,鼓励双侧受力平衡,体现人类运动控制原则。在TacSL基准测试中,带对称性正则化的CMT实现96.59%的插接成功率,优于朴素融合与门控融合基线,并接近特权配置(腕部+接触力,96.09%)。结果表明:(i)触觉感知对精确定位不可或缺;(ii)基于物理先验的多模态融合能充分释放视觉与触觉的互补优势,实现在真实传感条件下的近特权性能。
原文摘要 · Abstract (English)
Insertion tasks in robotic manipulation demand precise, contact-rich interactions that vision alone cannot resolve. While tactile feedback is intuitively valuable, existing studies have shown that naïve visuo-tactile fusion often fails to deliver consistent improvements. In this work, we propose a Cross-Modal Transformer (CMT) for visuo-tactile fusion that integrates wrist-camera observations with tactile signals through structured self- and cross-attention. To stabilize tactile embeddings, we further introduce a physics-informed regularization that encourages bilateral force balance, reflecting principles of human motor control. Experiments on the TacSL benchmark show that CMT with symmetry regularization achieves a 96.59% insertion success rate, surpassing naïve and gated fusion baselines and closely matching the privileged "wrist + contact force" configuration (96.09%). These results highlight two central insights: (i) tactile sensing is indispensable for precise alignment, and (ii) principled multimodal fusion, further strengthened by physics-informed regularization, unlocks complementary strengths of vision and touch, approaching privileged performance under realistic sensing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。