arXiv:2507.13318cs.CL2025-07EMNLP被引 6

构建首个振动触觉图文数据集,助力理解用户对震动的感知体验。

HapticCap: A Multimodal Dataset and Task for Understanding User Experience of Vibration Haptic Signals

  • 构建9.2万组触觉信号与人工描述配对的数据集,覆盖感官、情感与联想属性。
  • 提出触觉描述检索任务,融合T5与AST模型在分类别训练下表现最佳。
  • 解决触觉描述稀缺与文本生成能力弱的难题,适合人机交互与虚拟现实研究者。

触觉信号(如手机震动、虚拟现实触感反馈)能有效传递信息并增强真实感,但设计出能引起用户共鸣的信号仍具挑战。为此,我们引入首个完全人工标注的多模态触觉-文本数据集HapticCap,包含92,070组触觉信号与用户对感官、情感及联想属性的文字描述。基于此,我们提出触觉描述检索任务,并采用监督对比学习框架,将同类别文本与振动信号表征对齐。实验表明,语言模型T5与音频模型AST的结合在该任务中表现最优,尤其在针对不同描述类别分别训练时效果更佳。

原文摘要 · Abstract (English)

Haptic signals, from smartphone vibrations to virtual reality touch feedback, can effectively convey information and enhance realism, but designing signals that resonate meaningfully with users is challenging. To facilitate this, we introduce a multimodal dataset and task, of matching user descriptions to vibration haptic signals, and highlight two primary challenges: (1) lack of large haptic vibration datasets annotated with textual descriptions as collecting haptic descriptions is time-consuming, and (2) limited capability of existing tasks and models to describe vibration signals in text. To advance this area, we create HapticCap, the first fully human-annotated haptic-captioned dataset, containing 92,070 haptic-text pairs for user descriptions of sensory, emotional, and associative attributes of vibrations. Based on HapticCap, we propose the haptic-caption retrieval task and present the results of this task from a supervised contrastive learning framework that brings together text representations within specific categories and vibrations. Overall, the combination of language model T5 and audio model AST yields the best performance in the haptic-caption retrieval task, especially when separately trained for each description category.

触觉感知多模态数据集文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。