arXiv:2505.20032cs.CVcs.LG2025-05被引 5

提出视觉触觉位置编码,提升多模态融合与零样本泛化能力

ViTaPEs: Visuotactile Position Encodings for Cross-Modal Alignment in Multimodal Transformers

  • 分两阶段注入位置编码:局部编码保模态特征,全局编码促跨模态对齐
  • 在多个真实数据集上超越现有方法,实现未见场景的零样本迁移
  • 适合机器人抓取等需要多模态感知的任务研究者使用

触觉感知提供纹理、柔度和力等局部关键信息,与视觉感知互补。尽管视觉触觉表征学习取得进展,但模态融合与跨任务、跨环境泛化仍面临挑战,且多数方法未考虑位置编码,忽视了捕捉细粒度跨模态关联所需的多阶段空间推理。本文提出基于Transformer的ViTaPEs架构,从配对的视觉与触觉输入中学习任务无关的表征。核心思想是两阶段位置编码注入:在各模态流内添加局部(模态特定)位置编码,在联合标记序列前添加全局位置编码,于注意力计算前提供共享的位置词汇。明确位置注入点并进行受控消融,比较在标记级非线性前与自注意力前注入的效果。在多个大规模真实世界数据集上的实验表明,ViTaPEs不仅在多种识别任务中超越当前最优基线,还在未见过的域外场景中展现零样本泛化能力。进一步在机器人抓取任务中验证其迁移学习优势,预测抓取成功率优于现有最优方法。

原文摘要 · Abstract (English)

Tactile sensing provides local essential information that is complementary to visual perception, such as texture, compliance, and force. Despite recent advances in visuotactile representation learning, challenges remain in fusing these modalities and generalizing across tasks and environments without heavy reliance on pre-trained vision-language models. Moreover, existing methods do not study positional encodings, thereby overlooking the multi-stage spatial reasoning needed to capture fine-grained visuotactile correlations. We introduce ViTaPEs, a transformer-based architecture for learning task-agnostic visuotactile representations from paired vision and tactile inputs. Our key idea is a two-stage positional injection: local (modality-specific) positional encodings are added within each stream, and a global positional encoding is added on the joint token sequence immediately before attention, providing a shared positional vocabulary at the stage where cross-modal interaction occurs. We make the positional injection points explicit and conduct controlled ablations that isolate their effect before a token-wise nonlinearity versus immediately before self-attention. Experiments on multiple large-scale real-world datasets show that ViTaPEs not only surpasses state-of-the-art baselines across various recognition tasks but also demonstrates zero-shot generalization to unseen, out-of-domain scenarios. We further demonstrate the transfer-learning strength of \emph{ViTaPEs} in a robotic grasping task, where it outperforms state-of-the-art baselines in predicting grasp success. Project page: https://sites.google.com/view/vitapes

多模态融合触觉感知位置编码机器人抓取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。