用语言模型统一理解说话与动作,让虚拟人自然互动。
The Language of Motion: Unifying Verbal and Non-verbal Language of 3D Human Motion
- 构建多模态语言模型,输入可为文字、语音或动作
- 在共说话手势生成任务上达顶尖水平,且训练数据少
- 支持可编辑手势和从动作预测情绪,适合虚拟角色开发
人类交流本质是多模态的,包含语言、面部表情和肢体动作。建模这些行为对理解人际互动及创造自然沟通的虚拟角色至关重要,适用于游戏、影视和虚拟现实。然而,现有运动生成模型通常仅支持单一输入模态——如语音、文本或运动数据,无法充分利用多样数据。本文提出一种新框架,通过多模态语言模型统一处理语言与非语言信息,灵活接收文本、语音、动作或其组合作为输入。结合创新的预训练策略,该模型不仅在共说话手势生成任务上达到当前最佳性能,且训练所需数据量显著减少。此外,模型还实现一系列新任务,如可编辑的手势生成和从运动中预测情绪。我们认为,统一语言与动作表达对真实场景应用至关重要,而语言模型为此提供了强大路径。项目页面:languageofmotion.github.io。
原文摘要 · Abstract (English)
Human communication is inherently multimodal, involving a combination of verbal and non-verbal cues such as speech, facial expressions, and body gestures. Modeling these behaviors is essential for understanding human interaction and for creating virtual characters that can communicate naturally in applications like games, films, and virtual reality. However, existing motion generation models are typically limited to specific input modalities -- either speech, text, or motion data -- and cannot fully leverage the diversity of available data. In this paper, we propose a novel framework that unifies verbal and non-verbal language using multimodal language models for human motion understanding and generation. This model is flexible in taking text, speech, and motion or any combination of them as input. Coupled with our novel pre-training strategy, our model not only achieves state-of-the-art performance on co-speech gesture generation but also requires much less data for training. Our model also unlocks an array of novel tasks such as editable gesture generation and emotion prediction from motion. We believe unifying the verbal and non-verbal language of human motion is essential for real-world applications, and language models offer a powerful approach to achieving this goal. Project page: languageofmotion.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。