让机器人更灵敏地触觉反馈,实现实时精准操作。
AT-VLA: Adaptive Tactile Injection for Enhanced Feedback Reaction in Vision-Language-Action Models

- 动态决定触觉注入时机与位置,减少对预训练模型干扰。
- 触觉响应延迟仅0.04秒,实现毫秒级闭环控制。
- 适合需要精细触觉交互的机器人操作任务。
视觉-语言-动作(VLA)模型在执行多样化任务方面取得了显著进展,但在需要精确物理交互的接触密集型操作中仍面临挑战。为解决此问题,现有研究尝试在下游任务中引入触觉信号,使预训练的VLA能够理解触觉反馈。然而,在微调阶段引入预训练阶段罕见的新模态,可能破坏VLA的预训练能力。此外,VLA固有的慢推理速度限制了触觉反馈在动作调整中的有效利用。为此,我们提出自适应触觉注入的视觉-语言-动作模型(AT-VLA),其核心是自适应触觉注入机制,动态决定触觉信号注入的时间与位置,仅在显著提升动作生成时才引入,从而最小化对预训练表示的干扰。同时,提出触觉反应双流机制,将感知处理解耦为低频的视觉-语言流用于深层推理,和高频的触觉控制流用于快速物理交互理解,实现0.04秒内的实时闭环响应。真实世界实验充分验证了AT-VLA在接触密集型操作任务中的有效性。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have significantly advanced the capabilities of robotic agents in executing diverse tasks; however, they still face challenges in contact-rich manipulation scenarios that require precise physical interactions. To address this limitation, recent studies have attempted to incorporate tactile signals during downstream tasks, enabling pretrained VLAs to interpret tactile feedback. Nevertheless, introducing new modalities during finetuning, which are rarely present in the pretrain stage, may disrupt the pretrained capabilities of VLAs. In addition, the inherently slow inference speed of VLAs hampers real-time responsiveness and limits the effective utilization of tactile feedback for action adjustment. To overcome these challenges, we propose Adaptive Tactile Vision-Language-Action (AT-VLA), which introduces a novel Adaptive Tactile Injection mechanism. This mechanism dynamically determines the appropriate timing and locations for tactile injection, incorporating only when it significantly contributes to action generation, thereby minimizing interference with pretrained representations. Furthermore, to enable rapid and accurate tactile responses, we propose a Tactile Reaction Dual-Stream mechanism, which decouples sensory processing into a slow visual-language stream for low-frequency perceptual reasoning and a fast tactile control stream for high-frequency physical interaction understanding, achieving real-time close-loop responses within 0.04 s. Real-world experiments thoroughly validate the effectiveness of AT-VLA in contact-rich manipulation tasks. The project page is available at: https://sites.google.com/view/at-vla.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。