让机器人通过触觉理解接触状态,提升抓取与操作能力
CLTP: Contrastive Language-Tactile Pre-training for 3D Contact Geometry Understanding
- 用5万+触觉点云与语言配对数据训练,聚焦接触位置、形状和力度
- 在零样本分类和触觉交互任务中显著优于现有方法
- 适合做具身智能、机器人触觉感知与多模态模型研究者
将触觉感知与视觉-语言模型结合的进展展示了机器人多模态感知的巨大潜力。然而,现有触觉描述仍局限于纹理等表层属性,忽略了机器人操作中关键的接触状态。为此,我们提出CLTP,一种直观且高效的语言-触觉预训练框架,将触觉3D点云与自然语言在多种接触场景中对齐,实现对接触状态敏感的触觉语言理解,适用于接触丰富的操作任务。我们首次构建了一个包含5万+触觉3D点云-语言配对的新数据集,其描述从触觉传感器视角明确捕捉多维接触状态(如接触位置、形状和力)。CLTP利用预对齐且冻结的视觉-语言特征空间,连接整体文本与触觉模态。实验验证其在三个下游任务中的优越性:零样本3D分类、接触状态分类以及触觉3D大语言模型交互。据我们所知,这是首个从接触状态视角对齐触觉与语言表示以支持操作任务的研究,为触觉-语言-动作模型学习提供了巨大潜力。代码与数据集已开源。
原文摘要 · Abstract (English)
Recent advancements in integrating tactile sensing with vision-language models (VLMs) have demonstrated remarkable potential for robotic multimodal perception. However, existing tactile descriptions remain limited to superficial attributes like texture, neglecting critical contact states essential for robotic manipulation. To bridge this gap, we propose CLTP, an intuitive and effective language tactile pretraining framework that aligns tactile 3D point clouds with natural language in various contact scenarios, thus enabling contact-state-aware tactile language understanding for contact-rich manipulation tasks. We first collect a novel dataset of 50k+ tactile 3D point cloud-language pairs, where descriptions explicitly capture multidimensional contact states (e.g., contact location, shape, and force) from the tactile sensor's perspective. CLTP leverages a pre-aligned and frozen vision-language feature space to bridge holistic textual and tactile modalities. Experiments validate its superiority in three downstream tasks: zero-shot 3D classification, contact state classification, and tactile 3D large language model (LLM) interaction. To the best of our knowledge, this is the first study to align tactile and language representations from the contact state perspective for manipulation tasks, providing great potential for tactile-language-action model learning. Code and datasets are open-sourced at https://sites.google.com/view/cltp/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。