通过对齐触觉监督位置,提升机器人接触操作的感知与动作协同能力。
Representation-Aligned Tactile Grounding for Contact-Rich Robotic Manipulation

- 从中间动作特征预测未来触觉嵌入,实现触觉与动作的精准对齐。
- 在真实世界接触任务中,性能优于非对齐或多接口触觉预测方法。
- 适合研究触觉增强型机器人控制、具身智能与多模态学习的学者。
触觉增强的视觉-语言-动作(VLA)策略被用于接触丰富的操作任务,其中关键交互状态常隐藏于视觉之外。未来触觉预测是一种有前景的方法,它将触觉结果转化为动作引发接触动态的监督信号。然而,VLA策略包含不同功能的表征,从感知编码到运动预测,导致难以确定监督应施加的位置。本文将其视为表征对齐问题。通过线性探测分析发现,未来触觉状态最可预测于中间动作专家特征,而非视觉-语言特征或最终动作状态。受此启发,提出轻量级隐空间触觉预测器(LTP),从识别出的中间表示预测紧凑的未来触觉嵌入。通过避免直接预测噪声较大的原始触觉信号,LTP提供了与动作结果对齐的接地信号,使中间动作表征与未来接触后果相一致。在真实世界接触丰富操作任务上的实验表明,表征对齐的触觉接地优于低对齐或多接口触觉预测,凸显了触觉监督施加位置的重要性。
原文摘要 · Abstract (English)
Tactile-enhanced vision-language-action (VLA) policies have been introduced for contact-rich manipulation, where critical interaction states are often hidden from vision. Future tactile prediction is a promising way to use touch because it turns tactile outcomes into supervision for action-induced contact dynamics. Yet VLA policies contain representations with different roles, from perceptual encoding to motor prediction, making it unclear where this supervision should be applied. We study this as a representation-alignment problem. Through a linear probe analysis, we find that future tactile states are most predictable from intermediate action-expert features, rather than from vision-language features or final action states. Motivated by this observation, we introduce a lightweight Latent Tactile Predictor (LTP), which predicts compact future tactile embeddings from the identified intermediate representation. By avoiding direct prediction of noisy raw tactile signals, LTP provides an action-outcome grounding signal that aligns intermediate action representations with future contact consequences. Experiments on real-world contact-rich manipulation tasks show that representation-aligned tactile grounding outperforms less aligned or multi-interface tactile prediction, highlighting the importance of where tactile supervision is applied.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。