构建通用触觉模型,让设备更智能地感知和响应物理交互。
The Potential of Haptic Foundation Models

- 提出触觉基础模型新范式,融合动作与物理动态建模。
- 在多个数据集上验证模型在力估计、滑动检测上的高精度表现。
- 适合机器人、可穿戴设备等需要自适应触觉交互的领域。
尽管语言与视觉领域的基础模型已取得成功,但其在具身智能中的拓展受限于缺乏通用触觉感知能力。这一瓶颈在消费电子领域尤为突出:智能手机、可穿戴设备、VR控制器、家用机器人及健康监测设备均需安全且自适应的物理交互。受硬件异构性及主动物理数据采集需求制约,现有触觉模型仍高度任务特定。本文探讨触觉基础模型(Haptic Foundation Models, HFMs)的变革潜力与发展路径,强调从被动大语言模型和视觉语言模型向主动HFMs的范式转变,涵盖四个核心维度:动作耦合、物理动力学表示空间、连续时间序列数据粒度以及动作条件下的未来状态预测。此外,我们整合现有大规模触觉数据集,并在TacBench基准上对UniTouch、AnyTouch、T3和Sparsh进行力估计、滑动检测和相对位姿估计的评测。
原文摘要 · Abstract (English)
Despite the success of foundation models in language and vision, their expansion into embodied AI is bottlenecked by a lack of generalized touch sensing. This limitation is especially relevant to consumer electronics, where smartphones, wearables, VR controllers, home robots, and health monitoring devices require safe and adaptive physical interaction. Constrained by hardware heterogeneity and the necessity of active physical data collection, current haptic models remain rigidly task-specific. To overcome these limitations, this article explores the transformative potential and developmental trajectory of Haptic Foundation Models (HFMs). We detail the paradigm shift required to transition from passive Large Language Models and Vision Language Models into active HFMs across four core dimensions: action coupling, physical dynamical representation space, continuous time-series data granularity, and action-conditioned future state prediction. Furthermore, we synthesize existing large-scale tactile datasets and benchmark UniTouch, AnyTouch, T3, and Sparsh on TacBench for force estimation, slip detection, and relative pose estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。