arXiv:2505.22566cs.CVcs.AI2025-05NeurIPS被引 14

首个融合视觉与触觉的视频理解大模型,让机器像人一样感知物体物理属性。

Universal Visuo-Tactile Video Understanding for Embodied Interaction

  • 构建三阶段训练框架,实现视觉与触觉信息的深度融合。
  • 在15万帧数据上训练,可精准识别硬度、弹性等四大触觉属性。
  • 适合机器人交互、智能硬件研发人员,推动具身智能发展。

触觉感知对具身智能体理解无法仅通过视觉判断的物体物理属性至关重要。现有方法虽在视觉和语言模态上取得进展,却未能有效整合提供关键触觉反馈的触觉信息。本文提出VTV-LLM,首个面向通用视觉-触觉视频(VTV)理解的多模态大语言模型,弥合触觉感知与自然语言之间的鸿沟。为应对跨传感器、跨模态融合挑战,我们构建了包含100种不同物体、3种触觉传感器(GelSight Mini、DIGIT、Tac3D)采集的15万帧视频数据集VTV150K,标注了硬度、凸起、弹性、摩擦力四个基本触觉属性。提出新颖的三阶段训练范式:先进行VTV增强以获得鲁棒的视觉-触觉表示,再进行VTV-文本对齐以建立跨模态对应关系,最后通过文本提示微调实现自然语言生成。该框架支持特征评估、对比分析、场景化决策等复杂触觉推理能力。实验表明,VTV-LLM在触觉视频理解任务中表现优异,为更直观的人机触觉交互奠定了基础。

原文摘要 · Abstract (English)

Tactile perception is essential for embodied agents to understand physical attributes of objects that cannot be determined through visual inspection alone. While existing approaches have made progress in visual and language modalities for physical understanding, they fail to effectively incorporate tactile information that provides crucial haptic feedback for real-world interaction. In this paper, we present VTV-LLM, the first multi-modal large language model for universal Visuo-Tactile Video (VTV) understanding that bridges the gap between tactile perception and natural language. To address the challenges of cross-sensor and cross-modal integration, we contribute VTV150K, a comprehensive dataset comprising 150,000 video frames from 100 diverse objects captured across three different tactile sensors (GelSight Mini, DIGIT, and Tac3D), annotated with four fundamental tactile attributes (hardness, protrusion, elasticity, and friction). We develop a novel three-stage training paradigm that includes VTV enhancement for robust visuo-tactile representation, VTV-text alignment for cross-modal correspondence, and text prompt finetuning for natural language generation. Our framework enables sophisticated tactile reasoning capabilities including feature assessment, comparative analysis, scenario-based decision making and so on. Experimental evaluations demonstrate that VTV-LLM achieves superior performance in tactile video understanding tasks, establishing a foundation for more intuitive human-machine interaction in tactile domains.

具身智能多模态触觉感知大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。