通过触觉与声音融合,提升机器人识别人类情绪与社交手势的能力
Touch and Tell: Multimodal Decoding of Human Emotions and Social Gestures for Robots
- 在机器人上集成压力传感器与麦克风,捕捉触觉与声音信号
- 多模态模型对10种情绪分类准确率达40%,6类社交手势达90.74%
- 触觉+声音融合显著优于单一模态,适合人机情感交互研究
人类情绪可通过细微的触觉动作传递。以往研究多关注人类如何通过触觉感知情绪,或为机器人提取情绪特征,但缺乏对情绪与手势通过触觉可靠传达并由数据驱动方法解析的理解。本研究通过在社交机器人上集成自定义压阻式压力传感器与麦克风,探索触觉与声音在情绪和手势表达中的一致性与可区分性。28名参与者先自发用触觉表达10种情绪,再执行6种预设社交触觉动作。结果表明,情绪与动作表达具统计显著一致性,但部分情绪的组内相关性较低,相似唤醒度或效价的情绪难以区分。为此,我们构建了单模态与多模态模型(结合触觉与音频特征)。基于支持向量机(SVM)的多模态模型在10种情绪分类中达到40%准确率;采用卷积神经网络-长短期记忆网络(CNN-LSTM)的模型在手势分类中达90.74%准确率。结果表明,尽管单模态模型具备解码潜力,但触觉与声音融合显著提升情绪与手势解码性能。
原文摘要 · Abstract (English)
Human emotions are complex and can be conveyed through nuanced touch gestures. Previous research has primarily focused on how humans recognize emotions through touch or on identifying key features of emotional expression for robots. However, there is a gap in understanding how reliably these emotions and gestures can be communicated to robots via touch and interpreted using data driven methods. This study investigates the consistency and distinguishability of emotional and gestural expressions through touch and sound. To this end, we integrated a custom piezoresistive pressure sensor as well as a microphone on a social robot. Twenty-eight participants first conveyed ten different emotions to the robot using spontaneous touch gestures, then they performed six predefined social touch gestures. Our findings reveal statistically significant consistency in both emotion and gesture expression among participants. However, some emotions exhibited low intraclass correlation values, and certain emotions with similar levels of arousal or valence did not show significant differences in their conveyance. To investigate emotion and social gesture decoding within affective human-robot tactile interaction, we developed single-modality models and multimodal models integrating tactile and auditory features. A support vector machine (SVM) model trained on multimodal features achieved the highest accuracy for classifying ten emotions, reaching 40 %.For gesture classification, a Convolutional Neural Network- Long Short-Term Memory Network (CNN-LSTM) achieved 90.74 % accuracy. Our results demonstrate that even though the unimodal models have the potential to decode emotions and touch gestures, the multimodal integration of touch and sound significantly outperforms unimodal approaches, enhancing the decoding of both emotions and gestures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。