arXiv:2605.17336cs.ROcs.CV2026-05综述被引 3

系统梳理触觉与视觉语言融合的研究进展,构建分类框架。

Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms

论文配图:Tactile-based Multimodal Fusion in Embodied Intelligence: A Survey of Vision, Language, and Contact-Driven Paradigms
图 1 · 摘自论文原文
  • 按数据与方法双维度构建触觉多模态融合分类体系
  • 涵盖触觉-视觉、触觉-语言等多类数据集与交互任务
  • 适合机器人感知、具身智能研究者参考

触觉感知是具身智能的基础模态,能提供远程传感器无法替代的接触几何、材料特性与交互动态的直接反馈。然而,单一触觉感知受限于稀疏的空间覆盖和缺乏全局语义上下文。随着深度学习与大语言模型的发展,将触觉与视觉、语言融合已成为连接物理交互与语义推理的关键,催生了多模态触觉融合研究。尽管进展迅速,现有研究仍分散在不同数据集、传感模态与任务中,缺乏统一理论框架。本文综述截至2026年第一季度的多模态触觉融合研究,提出层次化分类体系,从数据侧分为触觉-视觉、触觉-语言、触觉-视觉-语言及触觉-视觉-其他数据集;从方法侧分为三大支柱:多模态感知与识别(聚焦物体理解与抓取预测)、跨模态生成(实现触觉、视觉、文本间的双向转换)、多模态交互(强调反馈控制与语言引导操作)。同时总结代表性触觉硬件、常用评估指标与基准设置,并讨论当前挑战与未来方向。

原文摘要 · Abstract (English)

Tactile sensing is a fundamental modality for embodied intelligence, offering unique and direct feedback on contact geometry, material properties, and interaction dynamics that remote sensors cannot replace. However, unimodal tactile perception is inherently limited by its sparse spatial coverage and lack of global semantic context. With the recent explosion in deep learning and large language models, integrating tactile with vision and language has become essential to bridge physical interaction with semantic reasoning, leading to the emergence of Multimodal Tactile Fusion. Despite rapid progress, the existing researches remain fragmented across disparate datasets, sensing modalities, and tasks, lacking a unified theoretical framework. To address this gap, this paper provides a comprehensive survey of multimodal tactile fusion research up to the first quarter of 2026. We propose a hierarchical taxonomy that organizes the field into two primary dimensions: multimodal datasets and multimodal methods. On the data side, we categorize resources ranging from Tactile-Vision datasets, Tactile-Language datasets, Tactile-Vision-Language datasets, and Tactile-Vision-Other datasets. On the method side, we structure prior work into three core pillars: (1) Multimodal Perception and Recognition, which focuses on object understanding and grasp prediction; (2) Cross-Modal Generation, focusing on bidirectional translation between tactile, vision, and text; and (3) Multimodal Interaction, emphasizing feedback control and language-guided manipulation. Furthermore, we summarize representative tactile sensing hardware, review commonly used evaluation metrics and benchmark settings, and discuss current challenges and promising future directions.

触觉融合具身智能多模态机器人感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。