arXiv:2603.12665cs.RO2026-03被引 15

让机器人通过触觉感知更精准地完成精细操作

TacVLA: Contact-Aware Tactile Fusion for Robust Vision-Language-Action Manipulation

  • 引入触觉感知的自适应融合机制,仅在接触时激活触觉信息
  • 在拆解任务中成功率提升20%,箱内抓取提升60%,遮挡下性能提高2.1倍
  • 适合需要高精度触觉反馈的机器人操作场景

视觉-语言-动作(VLA)模型在机器人操作中表现出显著优势,但其对视觉和语言的依赖导致在存在视觉遮挡、精细操作和物理接触的任务中表现不佳。为此,我们提出TacVLA,通过在基于Transformer的策略中融合触觉模态,增强精细操作能力。具体地,引入接触感知门控机制,仅在检测到接触时激活触觉令牌,实现自适应多模态融合,避免无关触觉干扰。融合后的视觉、语言和触觉令牌在Transformer架构中联合处理,强化接触密集交互中的跨模态对齐。在约束锁拆解、箱内抓取及鲁棒性评估中的大量实验表明,TacVLA优于基线模型,包括现有VLA模型和扩散策略,在拆解任务中平均成功率提升20%,箱内抓取提升60%,在视觉遮挡下性能提升2.1倍,并具备应对人为干扰的恢复能力。视频展示见 https://sites.google.com/view/tacvla。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have demonstrated significant advantages in robotic manipulation. However, their reliance on vision and language often leads to suboptimal performance in tasks involving visual occlusion, fine-grained manipulation, and physical contact. To address these challenges, we propose TacVLA, a fine-tuned VLA model by incorporating tactile modalities into the transformer-based policy to enhance fine-grained manipulation capabilities. Specifically, we introduce a contact-aware gating mechanism that selectively activates tactile tokens only when contact is detected, enabling adaptive multimodal fusion while avoiding irrelevant tactile interference. The fused visual, language, and tactile tokens are jointly processed within the transformer architecture to strengthen cross-modal grounding during contact-rich interaction. Extensive experiments on constraint-locked disassembly, in-box picking and robustness evaluations demonstrate that TacVLA outperforms baselines, %including existing VLA models and diffusion policies, improving the performance by averaging 20\% success rate in disassembly and 60\% in in-box picking, achieving a 2.1$\times$ improvement under visual occlusion, and showing recovery behavior under human disturbance. Videos are available at https://sites.google.com/view/tacvla.

机器人操作触觉融合多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。