arXiv:2502.12191cs.LGcs.CV2025-02ICLR被引 70

统一多传感器触觉表征,提升机器人感知与跨设备迁移能力

AnyTouch: Learning Unified Static-Dynamic Representation across Multiple Visuo-tactile Sensors

  • 构建多模态多传感器数据集TacQuad,支持静态与动态触觉融合
  • 提出AnyTouch框架,实现跨传感器的静态-动态统一表征学习
  • 在真实倒液体任务中表现优异,适用于多类型触觉传感器系统

视觉-触觉传感器旨在模拟人类触觉感知,使机器人能精确理解并操作物体。尽管已有多种精心设计的视觉-触觉传感器被集成到机器人系统中,但其数据特征差异大、标准化程度低,限制了强大触觉感知系统的建立。本文认为关键在于学习统一的多传感器表示,以实现传感器融合与触觉知识迁移。为此,我们引入TacQuad——一个来自四种不同视觉-触觉传感器的对齐多模态多传感器触觉数据集,支持多传感器显式集成。受人类通过纹理、压力变化等多维触觉信息感知环境启发,我们提出从静态和动态两个角度学习统一的多传感器表征。通过融合触觉图像与视频,我们构建AnyTouch框架,采用多层次结构,既捕捉像素级细节(通过掩码建模),又通过多模态对齐与跨传感器匹配学习语义级无传感器依赖特征,增强感知能力与迁移性。我们在多个数据集及真实世界的倒液体任务中验证方法,实验表明,该方法在各类传感器上均优于现有方法,具备出色的静态与动态感知能力。

原文摘要 · Abstract (English)

Visuo-tactile sensors aim to emulate human tactile perception, enabling robots to precisely understand and manipulate objects. Over time, numerous meticulously designed visuo-tactile sensors have been integrated into robotic systems, aiding in completing various tasks. However, the distinct data characteristics of these low-standardized visuo-tactile sensors hinder the establishment of a powerful tactile perception system. We consider that the key to addressing this issue lies in learning unified multi-sensor representations, thereby integrating the sensors and promoting tactile knowledge transfer between them. To achieve unified representation of this nature, we introduce TacQuad, an aligned multi-modal multi-sensor tactile dataset from four different visuo-tactile sensors, which enables the explicit integration of various sensors. Recognizing that humans perceive the physical environment by acquiring diverse tactile information such as texture and pressure changes, we further propose to learn unified multi-sensor representations from both static and dynamic perspectives. By integrating tactile images and videos, we present AnyTouch, a unified static-dynamic multi-sensor representation learning framework with a multi-level structure, aimed at both enhancing comprehensive perceptual abilities and enabling effective cross-sensor transfer. This multi-level architecture captures pixel-level details from tactile data via masked modeling and enhances perception and transferability by learning semantic-level sensor-agnostic features through multi-modal alignment and cross-sensor matching. We provide a comprehensive analysis of multi-sensor transferability, and validate our method on various datasets and in the real-world pouring task. Experimental results show that our method outperforms existing methods, exhibits outstanding static and dynamic perception capabilities across various sensors.

触觉感知多传感器融合表征学习机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。