arXiv:2501.00038cs.HCcs.RO2025-01被引 4

用触摸声音识别情绪与手势,保护隐私且高效。

Sound-Based Recognition of Touch Gestures and Emotions for Enhanced Human-Robot Interaction

  • 仅用触碰时的声音信号,不依赖视觉或触觉传感器。
  • 0.24M参数小模型,在不同音频长度下准确识别情绪与手势。
  • 适合注重隐私和低延迟的机器人交互场景。

情感识别与触觉手势解码对提升人机交互(HRI)至关重要,尤其在社交环境中,情感线索与触觉感知作用显著。然而,诸如Pepper、Nao和Furhat等类人机器人缺乏全身触觉皮肤,限制了其触觉交互能力。此外,基于视觉的情感识别常面临严格的GDPR合规挑战,因需收集个人面部数据。为解决上述局限并规避隐私问题,本文研究利用交互中触碰产生的声音来识别触觉手势,并分类情绪的唤醒度与效价维度。基于28名参与者与机器人Pepper的触觉互动数据集,设计了一种仅0.24M参数、0.94MB模型大小、0.7G FLOPs的轻量级纯音频识别模型。实验表明,该模型在不同输入音频长度下,均能有效识别情绪的唤醒度与效价状态,以及多种触觉手势。其性能接近知名预训练音频神经网络(PANNs),但计算开销大幅降低。

原文摘要 · Abstract (English)

Emotion recognition and touch gesture decoding are crucial for advancing human-robot interaction (HRI), especially in social environments where emotional cues and tactile perception play important roles. However, many humanoid robots, such as Pepper, Nao, and Furhat, lack full-body tactile skin, limiting their ability to engage in touch-based emotional and gesture interactions. In addition, vision-based emotion recognition methods usually face strict GDPR compliance challenges due to the need to collect personal facial data. To address these limitations and avoid privacy issues, this paper studies the potential of using the sounds produced by touching during HRI to recognise tactile gestures and classify emotions along the arousal and valence dimensions. Using a dataset of tactile gestures and emotional interactions from 28 participants with the humanoid robot Pepper, we design an audio-only lightweight touch gesture and emotion recognition model with only 0.24M parameters, 0.94MB model size, and 0.7G FLOPs. Experimental results show that the proposed sound-based touch gesture and emotion recognition model effectively recognises the arousal and valence states of different emotions, as well as various tactile gestures, when the input audio length varies. The proposed model is low-latency and achieves similar results as well-known pretrained audio neural networks (PANNs), but with much smaller FLOPs, parameters, and model size.

人机交互声音识别情绪识别隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。