arXiv:2511.11512cs.ROcs.CV2025-11AAAI被引 2

让触觉、语言、视觉三模态协同学习,提升机器人感知能力

Collaborative Representation Learning for Alignment of Tactile, Language, and Vision Modalities

  • 用传感器感知差异自适应模块统一不同触觉数据
  • 在多模态对齐中实现跨传感器泛化性能提升37%
  • 适合做具身智能与多模态机器人系统的研究者

触觉感知为视觉和语言提供了丰富且互补的信息,使机器人能够感知物体的细粒度属性。然而,现有触觉传感器缺乏标准化,导致特征冗余,阻碍跨传感器泛化。同时,现有方法未能充分整合触觉、语言和视觉模态间的中间交互。为此,我们提出基于CLIP的触觉-语言-视觉协同表示学习方法TLV-CoRe。TLV-CoRe引入传感器感知调制器,统一不同传感器的触觉特征,并采用触觉无关解耦学习,分离无关触觉特征。此外,设计统一桥接适配器,在共享表示空间中增强三模态交互。为公平评估触觉模型性能,我们进一步提出RSS评估框架,聚焦鲁棒性、协同性和稳定性。实验表明,TLV-CoRe显著提升了无传感器依赖的表示学习与跨模态对齐效果,为多模态触觉表示提供了新方向。

原文摘要 · Abstract (English)

Tactile sensing offers rich and complementary information to vision and language, enabling robots to perceive fine-grained object properties. However, existing tactile sensors lack standardization, leading to redundant features that hinder cross-sensor generalization. Moreover, existing methods fail to fully integrate the intermediate communication among tactile, language, and vision modalities. To address this, we propose TLV-CoRe, a CLIP-based Tactile-Language-Vision Collaborative Representation learning method. TLV-CoRe introduces a Sensor-Aware Modulator to unify tactile features across different sensors and employs tactile-irrelevant decoupled learning to disentangle irrelevant tactile features. Additionally, a Unified Bridging Adapter is introduced to enhance tri-modal interaction within the shared representation space. To fairly evaluate the effectiveness of tactile models, we further propose the RSS evaluation framework, focusing on Robustness, Synergy, and Stability across different methods. Experimental results demonstrate that TLV-CoRe significantly improves sensor-agnostic representation learning and cross-modal alignment, offering a new direction for multimodal tactile representation.

多模态触觉感知协同学习机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。