用多模态网络识别人对机器人自述,提升社交机器人共情能力
A Multimodal Neural Network for Recognizing Subjective Self-Disclosure Towards Social Robots
- 构建基于情感识别的多模态注意力网络,融合视听信息
- 新损失函数使F1得分达0.83,较基线提升0.48
- 适合研究人机交互与社会机器人认知的学者
主观自述是人类社交互动的重要特征。尽管社会科学文献已对自述的特征与影响有较多研究,但针对计算系统准确建模主观自述的工作仍较少,尤其缺乏对人类与机器人互动中自述行为的建模。随着社交机器人需在各类社交场景中与人类协作并建立关系,这一问题愈发紧迫。本文提出一种基于情感识别模型的定制化多模态注意力网络,利用自收集的大规模自述视频语料库进行训练,并设计了一种新的尺度保持交叉熵损失函数,优化分类与回归任务。结果表明,采用该损失函数的最佳模型在测试集上达到F1分数0.83,较最优基线模型提升0.48,显著推进了社交机器人识别交互伙伴自述的能力,这对具备社会认知功能的机器人至关重要。
原文摘要 · Abstract (English)
Subjective self-disclosure is an important feature of human social interaction. While much has been done in the social and behavioural literature to characterise the features and consequences of subjective self-disclosure, little work has been done thus far to develop computational systems that are able to accurately model it. Even less work has been done that attempts to model specifically how human interactants self-disclose with robotic partners. It is becoming more pressing as we require social robots to work in conjunction with and establish relationships with humans in various social settings. In this paper, our aim is to develop a custom multimodal attention network based on models from the emotion recognition literature, training this model on a large self-collected self-disclosure video corpus, and constructing a new loss function, the scale preserving cross entropy loss, that improves upon both classification and regression versions of this problem. Our results show that the best performing model, trained with our novel loss function, achieves an F1 score of 0.83, an improvement of 0.48 from the best baseline model. This result makes significant headway in the aim of allowing social robots to pick up on an interaction partner's self-disclosures, an ability that will be essential in social robots with social cognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。