首个统一触觉理解与生成的多模态模型,实现跨传感器触觉感知与仿真。
UniTac: A Unified Multimodal Model for Cross-Sensor Tactile Understanding and Generation

- 将触觉过程建模为非接触到接触的转变,双层表征融合传感器与物体属性。
- 在多传感器数据集上达成触觉理解新纪录,生成信号贴近真实触感。
- 适合机器人触觉感知、人机交互等需要真实触觉模拟的研究者使用。
统一多模态模型(UMMs)在跨模态理解与生成方面展现出巨大潜力,但现有研究极少将其拓展至触觉领域,而触觉意义由物体语义与传感器配置共同决定。为此,我们提出UniTac,首个专为触觉理解与生成设计的统一多模态模型。UniTac将触觉过程建模为从非接触到接触的演化,通过双层表征捕捉传感器与物体间的物理交互,同时编码两者属性。针对触觉理解,引入物体属性描述与传感器识别两项任务,增强对物理信息及跨传感器信息的推理能力;针对触觉生成,设计两阶段训练流程(重建与对齐),并采用基于传感器先验的采样策略,模拟真实触觉接触。在大规模多传感器数据集上训练后,UniTac在触觉理解任务中达到当前最优性能,并能生成跨传感器的逼真触觉信号。
原文摘要 · Abstract (English)
Unified multimodal models (UMMs) have shown great promise in integrating understanding and generation across diverse modalities. However, existing research rarely extends this paradigm to the tactile domain, where both object-level semantics and sensor-level configurations jointly determine the meaning of touch. To address this gap, we propose UniTac, the first UMM designed for tactile understanding and generation. UniTac models the tactile process as a transition from non-contact to contact, capturing the physical interaction between sensors and objects through a dual-level representation that encodes both sensor and object attributes. For tactile understanding, UniTac introduces two tasks, object property description and sensor identification, to enhance reasoning over physical and cross-sensor information. For tactile generation, we design a two-stage training paradigm consisting of reconstruction and alignment, together with a sensor-prior-based sampling strategy that simulates realistic tactile contact. Trained on large-scale multi-sensor datasets, UniTac achieves state-of-the-art performance in tactile understanding and generates realistic tactile signals across sensors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。