arXiv:2603.14604cs.ROcs.CV2026-03中稿 · ECCV被引 3

轻量级融合触觉信号,提升机器人操作的精准与稳定

Tactile Modality Fusion for Vision-Language-Action Models

  • 用特征调制技术在训练后融合触觉信息
  • 插入和开抽屉任务成功率显著提升
  • 适合需要精细触觉交互的机器人应用

我们提出 TacFiLM,一种轻量级多模态融合方法,将视觉-触觉信号整合进视觉-语言-动作(VLA)模型。尽管当前 VLA 模型已具备良好的泛化性和语义理解能力,但主要依赖视觉感知。然而,仅靠视觉无法捕捉接触密集操作中的复杂交互动态,如接触力、表面摩擦、柔顺性与剪切力。现有触觉融合方法常通过拼接标记或大规模预训练增加复杂度,而行为模型对计算资源要求高,亟需轻量级融合策略。TacFiLM 采用训练后微调方式,利用特征逐元素线性调制(FiLM)将预训练触觉表征条件化到中间视觉特征上。在插入与开抽屉任务上的实验表明,该方法在分布内与分布外任务中均显著提升了成功率、任务完成时间、执行效率及力稳定性。结果验证了 TacFiLM 是一种有效且高效的触觉融合方法,可显著改善接触密集型操作行为。

原文摘要 · Abstract (English)

We propose TacFiLM, a lightweight modality-fusion approach that integrates visual-tactile signals into vision-language-action (VLA) models. While advances in VLAs have introduced robot policies that are both generalizable and semantically grounded, these models mainly rely on vision-based perception. Vision alone, however, cannot capture the complex interaction dynamics that occur during contact-rich manipulation, including contact forces, surface friction, compliance, and shear. While recent attempts to integrate tactile signals into VLA models often increase complexity through token concatenation or large-scale pretraining, the heavy computational demands of behaviour models necessitate lightweight fusion strategies. To address these challenges, TacFiLM outlines a post-training finetuning approach that conditions intermediate visual features on pretrained tactile representations using feature-wise linear modulation (FiLM). Experimental results on insertion and drawer opening tasks demonstrate consistent improvements in success rate, direct task performance, completion time, and force stability across both in-distribution and out-of-distribution tasks. Together, these results support our method as an effective approach to integrating tactile signals into VLA models, improving contact-rich manipulation behaviours. Project page: https://charliem7.github.io/projects/TacFilm/

触觉融合机器人控制多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。