arXiv:2409.05427cs.CV2024-09AAAI被引 18

用文字精准生成触觉数据,提升机器人感知真实感

TextToucher: Fine-Grained Text-to-Touch Generation

  • 通过文本描述分层建模物体与传感器级触觉信息
  • 在扩散变换器中融合双粒度文本条件,生成高质量触觉样本
  • 提出对比式图文预训练评估法,适合触觉智能研究者

触觉感知在多模态大模型与具身智能发展中至关重要。为以最低成本收集触觉数据,现有研究尝试通过视觉到触觉图像转换生成触觉数据,但相比文本模态,视觉驱动的方法难以准确刻画人类触觉感受。本文从物体级(触觉纹理、触觉形状)与传感器级(凝胶状态)两个粒度详细分析触觉图像特征,利用文本描述建模这些信息,提出细粒度文本到触觉生成方法TextToucher。具体地,采用多模态大语言模型生成物体级触觉文本描述,并使用可学习的文本提示表示传感器级信息;为更好引导生成过程,融合双粒度文本信息,在扩散变换器架构中探索多种双粒度文本条件方法。此外,提出对比式文本-触觉预训练(CTTP)评估指标,精确衡量文本驱动生成的触觉数据质量。大量实验验证了TextToucher的优越性。源代码将公开于\url{https://github.com/TtuHamg/TextToucher}。

原文摘要 · Abstract (English)

Tactile sensation plays a crucial role in the development of multi-modal large models and embodied intelligence. To collect tactile data with minimal cost as possible, a series of studies have attempted to generate tactile images by vision-to-touch image translation. However, compared to text modality, visual modality-driven tactile generation cannot accurately depict human tactile sensation. In this work, we analyze the characteristics of tactile images in detail from two granularities: object-level (tactile texture, tactile shape), and sensor-level (gel status). We model these granularities of information through text descriptions and propose a fine-grained Text-to-Touch generation method (TextToucher) to generate high-quality tactile samples. Specifically, we introduce a multimodal large language model to build the text sentences about object-level tactile information and employ a set of learnable text prompts to represent the sensor-level tactile information. To better guide the tactile generation process with the built text information, we fuse the dual grains of text information and explore various dual-grain text conditioning methods within the diffusion transformer architecture. Furthermore, we propose a Contrastive Text-Touch Pre-training (CTTP) metric to precisely evaluate the quality of text-driven generated tactile data. Extensive experiments demonstrate the superiority of our TextToucher method. The source codes will be available at \url{https://github.com/TtuHamg/TextToucher}.

触觉生成文本生成扩散模型具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。