让虚拟人说话时手势更自然,靠语音上下文精准生成动作。
Contextual Gesture: Co-Speech Gesture Video Generation through Context-aware Gesture Representation
- 用语音上下文建模手势模式,提升动作语义准确性。
- 生成视频与语音对齐,支持长序列且像素级真实。
- 适合虚拟人、交互系统开发,可编辑手势细节。
对话手势生成对打造逼真虚拟人和增强人机交互至关重要,需使手势与语音同步。尽管已有进展,现有方法仍难以准确识别音频中的节奏或语义触发信号,以生成具上下文关联的手势模式,并实现像素级真实感。为此,我们提出 Contextual Gesture 框架,包含三个创新组件:(1) 时序对齐的语音-手势同步机制,(2) 通过知识蒸馏将语音上下文融入动作模式表示的上下文手势分词方法,(3) 利用边缘连接关联手势关键点的结构感知优化模块。大量实验表明,该框架不仅能生成与语音高度对齐、视觉逼真的手势视频,还支持长序列生成与视频手势编辑,如图1所示。
原文摘要 · Abstract (English)
Co-speech gesture generation is crucial for creating lifelike avatars and enhancing human-computer interactions by synchronizing gestures with speech. Despite recent advancements, existing methods struggle with accurately identifying the rhythmic or semantic triggers from audio for generating contextualized gesture patterns and achieving pixel-level realism. To address these challenges, we introduce Contextual Gesture, a framework that improves co-speech gesture video generation through three innovative components: (1) a chronological speech-gesture alignment that temporally connects two modalities, (2) a contextualized gesture tokenization that incorporate speech context into motion pattern representation through distillation, and (3) a structure-aware refinement module that employs edge connection to link gesture keypoints to improve video generation. Our extensive experiments demonstrate that Contextual Gesture not only produces realistic and speech-aligned gesture videos but also supports long-sequence generation and video gesture editing applications, shown in Fig.1.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。