让机器人同时看、感、说、动,提升抓取成功率与动作流畅度。
OmniVTLA: Vision-Tactile-Language-Action Models with Semantic-Aligned Tactile Sensing
- 设计双路径触觉编码器,融合视觉与力觉传感器数据。
- 在56类物体上构建13.5万组多模态数据集,提升触觉表征能力。
- 实测抓取成功率最高达96.9%,动作更平滑,适合复杂操作场景。
当前视觉-语言-动作(VLA)模型主要依赖视觉语言基础,在机器人操作中展现出任务泛化潜力。然而,由于触觉传感器异构性及数据获取困难,现有模型严重忽视触觉感知,难以应对接触密集型任务。为此,本文提出OmniVTLA,一种融合触觉感知的新架构。首先,采用双路径触觉编码框架,利用预训练视觉变压器(ViT)和语义对齐触觉ViT(SA-ViT),增强对多种视觉与力觉触觉传感器的感知能力。其次,构建ObjTac数据集,涵盖56个物体、10类,采集13.5万组包含文本、视觉与触觉信息的样本,填补现有跨模态触觉数据空白。第三,基于该数据集训练语义对齐触觉编码器,学习统一触觉表示,作为OmniVTLA的良好初始化。真实世界实验表明,相较于最先进基线,使用夹爪时成功率达96.9%(提升21.9%),使用灵巧手时达100%(提升6.2%),任务完成时间显著缩短,轨迹更平滑。
原文摘要 · Abstract (English)
Recent vision-language-action (VLA) models build upon vision-language foundations, and have achieved promising results and exhibit the possibility of task generalization in robot manipulation. However, due to the heterogeneity of tactile sensors and the difficulty of acquiring tactile data, current VLA models significantly overlook the importance of tactile perception and fail in contact-rich tasks. To address this issue, this paper proposes OmniVTLA, a novel architecture involving tactile sensing. Specifically, our contributions are threefold. First, our OmniVTLA features a dual-path tactile encoder framework. This framework enhances tactile perception across diverse vision-based and force-based tactile sensors by using a pretrained vision transformer (ViT) and a semantically-aligned tactile ViT (SA-ViT). Second, we introduce ObjTac, a comprehensive force-based tactile dataset capturing textual, visual, and tactile information for 56 objects across 10 categories. With 135K tri-modal samples, ObjTac supplements existing visuo-tactile datasets. Third, leveraging this dataset, we train a semantically-aligned tactile encoder to learn a unified tactile representation, serving as a better initialization for OmniVTLA. Real-world experiments demonstrate substantial improvements over state-of-the-art VLA baselines, achieving 96.9% success rates with grippers, (21.9% higher over baseline) and 100% success rates with dexterous hands (6.2% higher over baseline) in pick-and-place tasks. Besides, OmniVTLA significantly reduces task completion time and generates smoother trajectories through tactile sensing compared to existing VLA. Our ObjTac dataset can be found at https://readerek.github.io/Objtac.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。