让机器人通过触觉-力反馈理解物理交互,实现精准抓握与操作。
TaF-VLA: Tactile-Force Alignment in Vision-Language-Action Models for Force-aware Manipulation
- 用触觉-力对齐替代传统触觉-视觉对齐,捕捉动态物理规律。
- 构建1000万+条同步数据集,实现触觉与6维力矩的精确关联。
- 适合需要精细力控的机器人任务,如装配、柔顺操作。
视觉-语言-动作(VLA)模型在机器人操作中表现出强大通用性,但因其主要依赖视觉模态,缺乏接触密集型任务所需的物理直觉,难以实现精确力调节与物理推理。现有方法常将触觉输入视为辅助视觉纹理,忽略表面形变与交互动力学之间的深层关联。为此,我们提出从触觉-视觉对齐到触觉-力对齐的范式转变。提出TaF-VLA框架,将高维触觉观测显式地锚定于物理交互力。为此,我们开发了自动化触觉-力数据采集装置,构建了包含超过1000万条同步触觉观测、6轴力/力矩及矩阵力图的TaF-Dataset。为对齐时序触觉观测与交互力,核心组件为触觉-力适配器(TaF-Adapter),其作为触觉传感器编码器,提取离散化潜在信息以编码触觉观测。该机制确保学习表征能捕捉依赖历史、抗噪声的物理动态,而非静态视觉纹理。最终将此力对齐编码器集成至VLA主干网络。大量真实世界实验表明,TaF-VLA策略在接触密集型任务上显著优于最先进的触觉-视觉对齐与纯视觉基线,验证了其通过跨模态物理推理实现鲁棒力感知操作的能力。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have recently emerged as powerful generalists for robotic manipulation. However, due to their predominant reliance on visual modalities, they fundamentally lack the physical intuition required for contact-rich tasks that require precise force regulation and physical reasoning. Existing attempts to incorporate vision-based tactile sensing into VLA models typically treat tactile inputs as auxiliary visual textures, thereby overlooking the underlying correlation between surface deformation and interaction dynamics. To bridge this gap, we propose a paradigm shift from tactile-vision alignment to tactile-force alignment. Here, we introduce TaF-VLA, a framework that explicitly grounds high-dimensional tactile observations in physical interaction forces. To facilitate this, we develop an automated tactile-force data acquisition device and curate the TaF-Dataset, comprising over 10 million synchronized tactile observations, 6-axis force/torque, and matrix force map. To align sequential tactile observations with interaction forces, the central component of our approach is the Tactile-Force Adapter (TaF-Adapter), a tactile sensor encoder that extracts discretized latent information for encoding tactile observations. This mechanism ensures that the learned representations capture history-dependent, noise-insensitive physical dynamics rather than static visual textures. Finally, we integrate this force-aligned encoder into a VLA backbone. Extensive real-world experiments demonstrate that TaF-VLA policy significantly outperforms state-of-the-art tactile-vision-aligned and vision-only baselines on contact-rich tasks, verifying its ability to achieve robust, force-aware manipulation through cross-modal physical reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。