arXiv:2409.13726cs.CLcs.AI2024-09被引 9

跨文化非语言行为影响对话参与度,该研究用多语言数据验证了这一点。

Multilingual Dyadic Interaction Corpus NoXi+J: Toward Understanding Asian-European Non-verbal Cultural Characteristics and their Influences on Engagement

  • 构建了包含中日英法德五国的双人互动数据集NoXi+J,扩展了原数据。
  • 发现不同文化背景下非语言特征(如点头、语音节奏)对参与度预测有显著差异。
  • 适用于跨文化人机交互、情感计算及多语言对话系统开发者。

非语言行为是理解对话动态与交流者情感状态的核心挑战。尽管心理学研究已表明非语言行为存在文化差异,但针对这些差异的计算分析仍有限。为更深入理解跨文化语言群体中的参与度与非语言行为,本研究开展多语言计算分析,探究非语言特征在参与度识别与预测中的作用。首先,将原包含法国、德国和英国参与者互动数据的NoXi数据集扩展至包含日语和汉语双人对话数据,形成改进版的NoXi+J数据集。随后,通过多种模式识别技术提取多模态非语言特征,包括语音声学、面部表情、回应性言语和手势。进一步统计分析了倾听行为与回应模式,识别出各语言特有的文化依赖特征及多语言共有的普遍特征,并将其与对话双方的参与度相关联。最后,分析了文化差异对训练用于五语言数据集的LSTM模型输入特征的影响。结合SHAP分析与迁移学习,证实了特定语言数据集中输入特征的重要性与所分析文化特征之间存在显著关联。

原文摘要 · Abstract (English)

Non-verbal behavior is a central challenge in understanding the dynamics of a conversation and the affective states between interlocutors arising from the interaction. Although psychological research has demonstrated that non-verbal behaviors vary across cultures, limited computational analysis has been conducted to clarify these differences and assess their impact on engagement recognition. To gain a greater understanding of engagement and non-verbal behaviors among a wide range of cultures and language spheres, in this study we conduct a multilingual computational analysis of non-verbal features and investigate their role in engagement and engagement prediction. To achieve this goal, we first expanded the NoXi dataset, which contains interaction data from participants living in France, Germany, and the United Kingdom, by collecting session data of dyadic conversations in Japanese and Chinese, resulting in the enhanced dataset NoXi+J. Next, we extracted multimodal non-verbal features, including speech acoustics, facial expressions, backchanneling and gestures, via various pattern recognition techniques and algorithms. Then, we conducted a statistical analysis of listening behaviors and backchannel patterns to identify culturally dependent and independent features in each language and common features among multiple languages. These features were also correlated with the engagement shown by the interlocutors. Finally, we analyzed the influence of cultural differences in the input features of LSTM models trained to predict engagement for five language datasets. A SHAP analysis combined with transfer learning confirmed a considerable correlation between the importance of input features for a language set and the significant cultural characteristics analyzed.

跨文化分析非语言行为参与度预测多语言数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。