arXiv:2409.09276cs.RO2024-09中稿 · IROS2024, project …被引 6

用视觉语言模型融合触觉数据,实现无需训练的物体识别

Visuo-Tactile Zero-Shot Object Recognition with Vision-Language Model

  • 通过物体名称将触觉数据转为文本描述,实现零样本学习
  • 在FoodReplica和Cube数据集上显著提升难辨物体识别准确率
  • 适合机器人触觉感知、跨模态识别等应用

触觉感知对区分外观相似物体至关重要。本文提出一种将触觉数据融入视觉语言模型(VLM)的方法,用于视觉-触觉零样本物体识别。该方法利用VLM的零样本能力,从触觉相似物体的名称中推断触觉属性。所提方法仅需在训练时为每段触觉序列标注物体名称,即可将触觉数据转换为文本描述,具有低训练成本和强适应性。在FoodReplica和Cube数据集上的实验表明,该方法在仅靠视觉难以区分的物体识别任务中表现优异。

原文摘要 · Abstract (English)

Tactile perception is vital, especially when distinguishing visually similar objects. We propose an approach to incorporate tactile data into a Vision-Language Model (VLM) for visuo-tactile zero-shot object recognition. Our approach leverages the zero-shot capability of VLMs to infer tactile properties from the names of tactilely similar objects. The proposed method translates tactile data into a textual description solely by annotating object names for each tactile sequence during training, making it adaptable to various contexts with low training costs. The proposed method was evaluated on the FoodReplica and Cube datasets, demonstrating its effectiveness in recognizing objects that are difficult to distinguish by vision alone.

多模态零样本触觉感知视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。