让机器人通过触觉动态理解并预测接触,提升精细操作的准确性和鲁棒性。
UniTacVLA: Unified Tactile Understanding and Prediction in Vision Language Action Models

- 将触觉信号视为动态交互线索,构建统一触觉表征空间。
- 在4类接触密集任务中,成功率和操作精度显著优于现有方法。
- 适合需要高精度触觉反馈的机器人灵巧操作场景。
视觉-语言-动作(VLA)模型在许多机器人操作任务中表现优异,但在接触密集的灵巧操作任务中仍受限。为克服这一局限,近期的视觉-触觉-语言-动作(VTLA)方法将触觉感知引入VLA模型以提供直接接触信息,但通常将触觉信号作为被动辅助输入,难以建模触觉语义与未来物理交互。为此,我们提出一种统一的触觉学习框架,将触觉信号视为动态交互线索,用于接触理解与预测。具体而言,我们构建统一的触觉潜在空间,并通过触觉思维链推理与粗到细的未来触觉预测,联合建模当前触觉状态与未来接触变化,从而形成状态感知且动态感知的触觉先验。基于该先验,我们设计了触觉-动作混合控制器,结合实时与预测触觉反馈,对低频动作块进行高频修正。在四类接触密集任务(调整、插入、擦拭、装配)上,无论在干净环境还是外部扰动条件下,实验均表明该方法在成功率、操作精度和接触鲁棒性方面均优于现有方法,验证了其在灵巧物理交互中的有效性。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models have achieved strong performance in many robotic manipulation tasks, yet remain limited in contact-rich dexterous manipulation. To overcome this limitation, recent vision-tactile-language-action (VTLA) methods incorporate tactile sensing into VLA models to provide direct contact information. However, they typically treat tactile signals as passive auxiliary inputs, making it difficult to model tactile semantics and future physical interactions. To this end, we propose a unified tactile learning framework for contact-rich manipulation that models tactile signals as dynamic interaction cues for both contact understanding and prediction. Specifically, we construct a unified tactile latent space and jointly model current tactile states and future contact changes through tactile chain-of-thought reasoning and coarse-to-fine future tactile prediction, thereby forming a state-aware and dynamics-aware tactile prior. Based on this prior, we introduce a tactile-action mixed controller that combines real-time and predicted tactile feedback to refine low-frequency action chunks with high-frequency corrections. Real-world experiments on four categories of contact-rich tasks, including adjustment, insertion, wiping, and assembly, under both clean and externally perturbed settings, show that our method improves success rate, manipulation accuracy, and contact robustness over existing methods, demonstrating its effectiveness in dexterous physical interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。