用生成模型从动作数据中隐式学习情绪表达,提升虚拟形象情感传达能力。
Generative Learning as a Tool to Improve Perception of Emotional Body Motion Expressions

- 基于Transformer的生成模型从日本演员动作数据中学习情绪表达。
- 机器识别准确率22.80%,人类评估准确率24.91%,表现优于随机水平。
- 可增强情绪识别、提取典型动作模式、合成情绪强度过渡,适合情感计算研究者。
情绪性身体动作是非语言交流的重要组成部分,对虚拟现实化身与社交机器人具有重要意义。尽管生成模型为情绪动作学习带来新机遇,但因情绪线索微妙、个体差异和文化差异,准确生成仍具挑战。本文利用49名日本演员的表演动作数据集,训练基于Transformer的生成模型,以13类离散情绪标签为条件生成表达性动作。通过两方面评估:(1) 使用LSTM分类器评估机器可识别性,识别准确率达22.80%;(2) 对日本被试进行人类感知实验,评估与人类情感理解的一致性,准确率为24.91%。此外,验证了生成模型在三个实际任务中的效用:增强情绪识别模型、提取情绪特异性动作模式、合成情绪强度间的平滑过渡。结果表明,隐式数据驱动的生成建模有望提升情感计算应用与情绪表达的理解。
原文摘要 · Abstract (English)
Emotional body motion expressions are an essential element of non-verbal communication. Effectively conveying these expressions through technology is of utmost importance, for example, with virtual reality avatars and in social robotics. Recent advances in generative models have opened new opportunities for advancing research on emotional body motion learning. However, generating accurate emotional expression representations is challenging, given the subtlety of emotional cues, individual variability, and cultural differences. We investigate whether a generative model can implicitly learn emotional body motions directly from culturally grounded motion-capture data, without explicit emotion-motion guidance. Using a dataset of emotional performances by 49 Japanese actors, we trained a Transformer-based generative model to generate expressive motions conditioned on 13 discrete emotion labels. We evaluate the generated motions from two perspectives: (1) an LSTM-based classifier to assess recognizability by machine observers, achieving a recognition accuracy of 22.80%, and (2) a human perception study with Japanese raters to assess alignment with human affective interpretations, yielding a recognition accuracy of 24.91%. Beyond these, we evaluate the utility of generative modeling for three practical tasks: augmenting emotion recognition models, extracting representative emotion-specific motion patterns, and synthesizing smooth transitions between emotion intensities. Our findings highlight the potential of implicit, data-driven generative modeling to enhance affective computing applications and our understanding of emotion expressions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。