让机器人通过触觉实时反应,提升精细操作能力。
T-Rex: Tactile-Reactive Dexterous Manipulation

- 设计可变频率混合变压器架构,融合高频触觉信号。
- 在12个任务中成功率超基线30%以上,操控柔物更精准。
- 构建100小时触觉丰富数据集,支持复杂动作学习。
动态响应触觉信号是实现人类级灵巧操作的关键。然而,当前基于视觉-语言-动作(VLA)的机器人操控模型普遍忽略触觉模态,或仅限于静态触觉编码器,受限于多样化训练数据稀缺、评估标准缺失、现有VLA架构约束及静态触觉编码器能力不足。本文系统性突破上述瓶颈:提出一个大规模、100小时的触觉丰富数据集,采用新型高效采集方法聚焦基础运动原语;设计可变率混合变压器(MoT)架构,配备新颖的时间序列触觉VQ-VAE编码器,有效利用高频率触觉信号而不损失原有视觉-语言-动作能力;在需精细力控与柔物操作的12项任务中验证,触觉反应策略平均成功率显著优于最强基线,提升超30%。
原文摘要 · Abstract (English)
The ability to react dynamically to tactile signals has long been considered crucial to agile human-level dexterity. Yet contemporary learning-based Vision-Language-Action (VLA) models for robotic manipulation generally either overlook the tactile modality or are limited to encoders with static cues, due in part to the scarcity of diverse training data and standardized evaluation, architectural constraints in current VLA models, and limitations of static tactile encoders. In this paper, we push the frontier of tactile-reactive manipulation by addressing all of these limitations. We propose a large-scale, 100-hour tactile-rich dataset collected via a novel, data-efficient recipe that prioritizes elementary motor primitives. To effectively exploit naturally high-frequency touch signals without sacrificing the existing capabilities of existing VLAs, we introduce a variable-rate Mixture-of-Transformers (MoT) architecture equipped with a novel temporal tactile VQ-VAE encoder. We demonstrate the effectiveness of tactile-reactive policies on 12 manipulation tasks requiring delicate force control and deformable object manipulation, achieving over 30% higher average success rate than the strongest baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。