用大模型生成自然且可编辑的语音同步手势动画
LLM Gesticulator: Leveraging Large Language Models for Scalable and Controllable Co-Speech Gesture Synthesis
- 基于大语言模型构建语音驱动的手势生成框架
- 模型规模越大,动作质量越高,符合扩展规律
- 通过文本提示控制手势内容与风格,适合交互应用
本文提出 LLM Gesticulator,一个基于大语言模型的语音驱动式协同手势生成框架,可生成与输入音频节奏同步、动作自然且可编辑的全身动画。相比已有方法,该框架具备显著可扩展性:随着骨干大模型规模增大,评估指标呈现比例提升(即符合扩展规律)。同时,模型表现出强可控性,可通过文本提示调节生成手势的内容与风格。据我们所知,这是首个将大语言模型应用于协同手势生成任务的工作。通过现有客观指标与用户研究评估,本框架在生成质量上优于先前方法。
原文摘要 · Abstract (English)
In this work, we present LLM Gesticulator, an LLM-based audio-driven co-speech gesture generation framework that synthesizes full-body animations that are rhythmically aligned with the input audio while exhibiting natural movements and editability. Compared to previous work, our model demonstrates substantial scalability. As the size of the backbone LLM model increases, our framework shows proportional improvements in evaluation metrics (a.k.a. scaling law). Our method also exhibits strong controllability where the content, style of the generated gestures can be controlled by text prompt. To the best of our knowledge, LLM gesticulator is the first work that use LLM on the co-speech generation task. Evaluation with existing objective metrics and user studies indicate that our framework outperforms prior works.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。