用对比学习融合动作与语音,让计算机理解对话中的手势含义。
Learning Co-Speech Gesture Representations in Dialogue through Contrastive Learning: An Intrinsic Evaluation
- 通过多模态对比学习,从动作和语音中联合训练手势表征。
- 学习到的表征与人工标注的手势相似性高度相关,相关系数达0.67。
- 可恢复手势形态特征,适合做手势分析与人机交互研究。
在面对面对话中,共言语手势的形式与意义关系受上下文因素(如手势所指内容、说话者个体差异)影响而变化,这使得手势表征学习极具挑战。本文提出一种自监督对比学习方法,结合骨骼数据与语音信息,学习手势表征。采用包含大量具象手势的面对面对话数据集进行训练,并通过与人工标注的手势配对相似性对比,开展内在评估。此外,还进行了诊断探测分析,检验能否从表征中恢复可解释的特征。结果表明,学习到的表征与人类标注的相似性呈显著正相关(r=0.67),且其相似性模式与对话互动动态一致。同时,多个关于手势形态的特征可从潜在表征中成功恢复。研究证实,多模态对比学习是学习有意义手势表征的有前景方法,为大规模手势分析研究铺平了道路。
原文摘要 · Abstract (English)
In face-to-face dialogues, the form-meaning relationship of co-speech gestures varies depending on contextual factors such as what the gestures refer to and the individual characteristics of speakers. These factors make co-speech gesture representation learning challenging. How can we learn meaningful gestures representations considering gestures' variability and relationship with speech? This paper tackles this challenge by employing self-supervised contrastive learning techniques to learn gesture representations from skeletal and speech information. We propose an approach that includes both unimodal and multimodal pre-training to ground gesture representations in co-occurring speech. For training, we utilize a face-to-face dialogue dataset rich with representational iconic gestures. We conduct thorough intrinsic evaluations of the learned representations through comparison with human-annotated pairwise gesture similarity. Moreover, we perform a diagnostic probing analysis to assess the possibility of recovering interpretable gesture features from the learned representations. Our results show a significant positive correlation with human-annotated gesture similarity and reveal that the similarity between the learned representations is consistent with well-motivated patterns related to the dynamics of dialogue interaction. Moreover, our findings demonstrate that several features concerning the form of gestures can be recovered from the latent representations. Overall, this study shows that multimodal contrastive learning is a promising approach for learning gesture representations, which opens the door to using such representations in larger-scale gesture analysis studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。