arXiv:2604.03322cs.CVcs.AI2026-04被引 1

让机器人通过视觉+触觉+语言判断材质缺陷,比纯视觉更准

VitaTouch: Property-Aware Vision-Tactile-Language Model for Robotic Quality Inspection in Manufacturing

  • 融合视觉、触觉与语言,用对比学习对齐多模态信息
  • 在186个物体上实现88.9%硬度识别准确率,材料描述语义相似度达0.9009
  • 适合工业质检场景,尤其对反光、遮挡环境下的缺陷检测

智能制造中的质量检测需识别材料固有属性(如硬度、粗糙度),而仅依赖视觉的方法易受遮挡和反光影响。本文提出VitaTouch,一种面向材料属性推断与自然语言描述的视觉-触觉-语言模型。该模型采用模态专用编码器与双Q-Former提取语言相关特征,并压缩为大语言模型前缀令牌。通过对比学习显式对齐各模态与文本,并构建包含186个物体、5.1万张图像及5413组人工验证指令-回答对的VitaSet数据集。在HCT和整体TVL基准上表现最优,且在SSVTP上保持竞争力。在VitaSet上,硬度识别准确率达88.89%,粗糙度达75.13%,描述召回率54.81%;材料描述任务峰值语义相似度达0.9009。基于LoRA微调后,2/3/5类缺陷识别准确率分别为100.0%、96.0%、92.0%,闭环识别与端到端分拣成功率均达94.0%,在100次实验室机器人测试中验证有效。

原文摘要 · Abstract (English)

Quality inspection in smart manufacturing requires identifying intrinsic material and surface properties beyond visible geometry, yet vision-only methods remain vulnerable to occlusion and reflection. We propose VitaTouch, a property-aware vision-tactile-language model for material-property inference and natural-language attribute description. VitaTouch uses modality-specific encoders and a dual Q-Former to extract language-relevant visual and tactile features, which are compressed into prefix tokens for a large language model. We align each modality with text and explicitly couple vision and touch through contrastive learning. We also construct VitaSet, a multimodal dataset with 186 objects, 52k images, and 5.1k human-verified instruction-answer pairs. VitaTouch achieves the best performance on HCT and the overall TVL benchmark, while remaining competitive on SSVTP. On VitaSet, it reaches 88.89% hardness accuracy, 75.13% roughness accuracy, and 54.81% descriptor recall; the material-description task further achieves a peak semantic similarity of 0.9009. With LoRA-based fine-tuning, VitaTouch attains 100.0%, 96.0%, and 92.0% accuracy for 2-, 3-, and 5-category defect recognition, respectively, and delivers 94.0% closed-loop recognition accuracy and 94.0% end-to-end sorting success in 100 laboratory robotic trials. More details are available at the project page: https://vitatouch.github.io/

机器人质检多模态触觉感知工业应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。