arXiv:2505.02569cs.ROcs.HC2025-05被引 2

用视觉语言模型实现材料识别与温感反馈,提升虚拟触觉体验

HapticVLM: VLM-Driven Texture Recognition Aimed at Intelligent Haptic Interaction

  • 结合卷积网络与视觉语言模型,实时识别物体材质与环境温度
  • 材料识别准确率达84.67%,温感估计误差在8°C内达86.7%
  • 适合虚拟现实、辅助技术中的多模态触觉交互研究

本文提出HapticVLM,一种融合视觉-语言推理与深度卷积网络的新型多模态系统,支持实时触觉反馈。该系统采用基于ConvNeXt的材料识别模块生成鲁棒视觉嵌入,实现物体材质的精准识别;同时利用先进视觉语言模型Qwen2-VL-2B-Instruct,通过环境线索推断周围温度。系统通过扬声器输出振动触觉反馈,并借助珀耳帖模块提供热感信号,弥合视觉感知与触觉体验之间的鸿沟。实验表明,在五种不同听觉-触觉模式下平均识别准确率为84.67%,基于容差评估方法(8°C容差)在15种场景中温度估计准确率达86.7%。尽管结果令人鼓舞,但当前研究受限于少数显著模式及有限受试者群体。未来工作将拓展触觉模式范围并扩大用户测试以进一步优化与验证系统性能。总体而言,HapticVLM为上下文感知的多模态触觉交互迈出了重要一步,具有虚拟现实与辅助技术应用潜力。

原文摘要 · Abstract (English)

This paper introduces HapticVLM, a novel multimodal system that integrates vision-language reasoning with deep convolutional networks to enable real-time haptic feedback. HapticVLM leverages a ConvNeXt-based material recognition module to generate robust visual embeddings for accurate identification of object materials, while a state-of-the-art Vision-Language Model (Qwen2-VL-2B-Instruct) infers ambient temperature from environmental cues. The system synthesizes tactile sensations by delivering vibrotactile feedback through speakers and thermal cues via a Peltier module, thereby bridging the gap between visual perception and tactile experience. Experimental evaluations demonstrate an average recognition accuracy of 84.67% across five distinct auditory-tactile patterns and a temperature estimation accuracy of 86.7% based on a tolerance-based evaluation method with an 8°C margin of error across 15 scenarios. Although promising, the current study is limited by the use of a small set of prominent patterns and a modest participant pool. Future work will focus on expanding the range of tactile patterns and increasing user studies to further refine and validate the system's performance. Overall, HapticVLM presents a significant step toward context-aware, multimodal haptic interaction with potential applications in virtual reality, and assistive technologies.

触觉交互多模态视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。