arXiv:2607.14728cs.CVcs.RO2026-07被引 2

用少量数据实现跨传感器触觉生成,提升机器人感知泛化能力

VQ-Touch: A Data-Efficient Tactile Generation Framework Across Sensors and Scenarios

论文配图:VQ-Touch: A Data-Efficient Tactile Generation Framework Across Sensors and Scenarios
图 1 · 摘自论文原文
  • 通过离散变分自编码器学习复杂触觉特征
  • 少样本混合训练使模型兼容主流触觉传感器
  • 支持图像与标签多模态生成,适合资源受限场景

触觉图像生成能显著降低对昂贵且易损传感器的依赖,通过合成高保真触觉数据,为机器人感知和人机交互系统提供高效的信息获取方案。然而,现有方法依赖特定传感器的大规模多样化数据集,存在数据利用效率低、泛化能力差的问题,在视觉受限环境下表现不佳。为此,我们提出VQ-Touch,一个支持跨传感器与多场景应用的触觉生成框架。具体而言,为高效提取数据中的复杂形变与纹理特征,我们设计了DM-VQGAN,一种高效的触觉表征学习器。此外,引入具有统一条件接口的离散扩散解码器,支持图像与标签等多模态生成任务,并通过少样本混合训练增强模型泛化能力,从而实现对当前主流触觉传感器及其变体的兼容。实验表明,VQ-Touch在多个任务中超越现有先进方法。

原文摘要 · Abstract (English)

Tactile image generation significantly reduces the dependency on expensive and wear-prone sensors by synthesizing high-fidelity tactile data, offering an efficient solution for tactile information acquisition in robotic perception and human-machine interaction systems. However, existing methods depend on large-scale, diverse datasets from specific sensors and lack efficient data utilization and robust generalization capabilities, struggling in vision-limited environments. To address this, we introduce VQ-Touch, a tactile generation framework that supports both cross-sensor and multi-scenario applications. Specifically, to efficiently extract complex deformation and texture features from the data, we propose DM-VQGAN, an effective tactile representation learner. Furthermore, we introduce a discrete diffusion decoder with a unified conditioning interface, supporting multimodal generation tasks such as images and labels, and enhances the model's generalization capability through few-shot mixed training, thus achieving compatibility with current mainstream sensors and their variants. Experiments show that VQ-Touch surpasses state-of-the-art methods in multiple tasks.

触觉生成少样本学习多模态机器人感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。