让机器人通过触觉理解力矩、接触形状等精细物理信息,提升抓取精准度。
FG-CLTP: Fine-Grained Contrastive Language Tactile Pretraining for Robotic Manipulation
- 用10万组触觉点云与语言配对数据,量化捕捉力、接触面等物理状态
- 触觉分类准确率达95.9%,回归误差比现有方法降低52.6%
- 适合做高精度触觉-语言-动作协同控制的机器人研究者
将触觉感知融入视觉-语言-动作(VLA)模型已展现出变革性潜力。然而,现有触觉表征多依赖纹理等定性描述,忽视力大小、接触几何、主轴方向等定量接触状态,这些对精细操作至关重要。为此,我们提出细粒度对比语言触觉预训练框架FG-CLTP。首先构建包含超10万组触觉3D点云-语言配对的新数据集,从传感器视角显式捕捉多维接触状态。随后引入离散化数值标记机制,实现定量与语义对齐,将物理量直接注入多模态特征空间。所提模型分类准确率达95.9%,回归误差(MAE)相比最先进方法降低52.6%。此外,3D点云表示实现传感器无关性,模拟到现实的差距仅3.5%。基于此细粒度表征,我们构建3D触觉-语言-动作(3D-TLA)架构,采用流匹配策略实现多模态推理与控制。大量实验表明,该框架在接触丰富的操作任务中显著优于强基线,为触觉-语言-动作模型提供稳健且可泛化的基础。
原文摘要 · Abstract (English)
Recent advancements in integrating tactile sensing into vision-language-action (VLA) models have demonstrated transformative potential for robotic perception. However, existing tactile representations predominantly rely on qualitative descriptors (e.g., texture), neglecting quantitative contact states such as force magnitude, contact geometry, and principal axis orientation, which are indispensable for fine-grained manipulation. To bridge this gap, we propose FG-CLTP, a fine-grained contrastive language tactile pretraining framework. We first introduce a novel dataset comprising over 100k tactile 3D point cloud-language pairs that explicitly capture multidimensional contact states from the sensor's perspective. We then implement a discretized numerical tokenization mechanism to achieve quantitative-semantic alignment, effectively injecting explicit physical metrics into the multimodal feature space. The proposed FG-CLTP model yields a 95.9% classification accuracy and reduces the regression error (MAE) by 52.6% compared to state-of-the-art methods. Furthermore, the integration of 3D point cloud representations establishes a sensor-agnostic foundation with a minimal sim-to-real gap of 3.5%. Building upon this fine-grained representation, we develop a 3D tactile-language-action (3D-TLA) architecture driven by a flow matching policy to enable multimodal reasoning and control. Extensive experiments demonstrate that our framework significantly outperforms strong baselines in contact-rich manipulation tasks, providing a robust and generalizable foundation for tactile-language-action models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。