用大模型让语音伴生手势精准模仿动作示例
Motion-example-controlled Co-speech Gesture Generation Leveraging Large Language Models
- 用大模型理解语音和动作示例,直接以示例为查询引导生成
- 在三个指标上达当前最好水平,保留原始动作细节
- 可控制身体各部位,支持视频、姿势、文本等多种输入
自动可控的语音伴生手势生成近年来受到广泛关注。现有方法通常通过预定义类别标签或从动作示例中隐式推导伪标签来实现控制,但常牺牲原始动作示例中的丰富细节。我们提出MECo框架,利用大语言模型(LLMs)进行微调,同时理解语音音频与动作示例,实现既能保留示例特异性又与语音一致的手势合成。不同于传统伪标签范式,我们将动作示例作为显式查询上下文嵌入提示结构中以指导生成。实验结果表明,该方法在三个指标上均达到当前最优:弗雷谢手势距离(FGD)、动作多样性与示例-手势相似性。此外,该框架支持对身体各部位的精细控制,并兼容多种输入模态,包括动作片段、静态姿态、人体视频序列及文本描述。代码、预训练模型与演示视频详见 https://robinwitch.github.io/MECo-Page。
原文摘要 · Abstract (English)
The automatic generation of controllable co-speech gestures has recently gained growing attention. While existing systems typically achieve gesture control through predefined categorical labels or implicit pseudo-labels derived from motion examples, these approaches often compromise the rich details present in the original motion examples. We present MECo, a framework for motion-example-controlled co-speech gesture generation by leveraging large language models (LLMs). Our method capitalizes on LLMs' comprehension capabilities through fine-tuning to simultaneously interpret speech audio and motion examples, enabling the synthesis of gestures that preserve example-specific characteristics while maintaining speech congruence. Departing from conventional pseudo-labeling paradigms, we position motion examples as explicit query contexts within the prompt structure to guide gesture generation. Experimental results demonstrate state-of-the-art performance across three metrics: Fréchet Gesture Distance (FGD), motion diversity, and example-gesture similarity. Furthermore, our framework enables granular control of individual body parts and accommodates diverse input modalities including motion clips, static poses, human video sequences, and textual descriptions. Our code, pre-trained models, and videos are available at https://robinwitch.github.io/MECo-Page.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。