让语音驱动的手势更自然,突出关键姿势的语义一致性。
Emphasizing Semantic Consistency of Salient Posture for Speech-Driven Gesture Generation
- 构建音频与姿态的联合表示空间,强化跨模态语义关联。
- 通过弱监督检测器识别关键姿势,重加权损失函数提升学习效果。
- 分枝生成面部与身体动作,提升细节表现力,适合影视动画应用。
语音驱动手势生成旨在合成与输入语音同步的手势序列。以往方法直接将紧凑的音频表示映射为手势序列,忽略了多模态间的语义关联,且难以处理关键手势。本文提出一种强调显著姿态语义一致性的新方法:首先学习音频与身体姿态的联合流形空间,挖掘两模态内在语义关联,并引入一致性损失进行约束;进一步通过弱监督检测器识别显著姿态,重新加权一致性损失,聚焦于显著姿态与语音高层语义的对应关系;此外,分别提取专用于面部表情和身体动作的音频特征,并设计独立分支实现面部与身体手势的合成。大量实验表明,该方法优于当前最先进方法。
原文摘要 · Abstract (English)
Speech-driven gesture generation aims at synthesizing a gesture sequence synchronized with the input speech signal. Previous methods leverage neural networks to directly map a compact audio representation to the gesture sequence, ignoring the semantic association of different modalities and failing to deal with salient gestures. In this paper, we propose a novel speech-driven gesture generation method by emphasizing the semantic consistency of salient posture. Specifically, we first learn a joint manifold space for the individual representation of audio and body pose to exploit the inherent semantic association between two modalities, and propose to enforce semantic consistency via a consistency loss. Furthermore, we emphasize the semantic consistency of salient postures by introducing a weakly-supervised detector to identify salient postures, and reweighting the consistency loss to focus more on learning the correspondence between salient postures and the high-level semantics of speech content. In addition, we propose to extract audio features dedicated to facial expression and body gesture separately, and design separate branches for face and body gesture synthesis. Extensive experimental results demonstrate the superiority of our method over the state-of-the-art approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。