用大模型生成多人舞蹈,保持音乐同步与舞者一致性。
Global Position Aware Group Choreography using Large Language Model
- 将多人舞蹈生成转为序列到序列翻译任务,利用大模型预测动作令牌。
- 生成舞蹈与音乐高度相关,且每位舞者动作连贯自然。
- 适合对舞蹈生成、多智能体协同感兴趣的开发者与研究者。
舞蹈是人类文化中深刻而普遍的表达形式,通过与音乐同步的动作传递情感与故事。尽管现有方法在单人舞蹈生成上已取得良好效果,但多人舞蹈生成仍属新兴领域。本文提出一种基于大语言模型(LLM)的群体编舞框架,将群体舞蹈生成建模为序列到序列的翻译任务。框架包含将连续特征转化为离散令牌的分词器,以及通过音频令牌预测动作令牌的微调大模型。通过合理的输入模态分词和精细的训练策略设计,该框架可生成真实且多样化的群体舞蹈,同时保持强音乐关联性与舞者间的一致性。大量实验与评估表明,本框架达到当前最优性能。
原文摘要 · Abstract (English)
Dance serves as a profound and universal expression of human culture, conveying emotions and stories through movements synchronized with music. Although some current works have achieved satisfactory results in the task of single-person dance generation, the field of multi-person dance generation remains relatively novel. In this work, we present a group choreography framework that leverages recent advancements in Large Language Models (LLM) by modeling the group dance generation problem as a sequence-to-sequence translation task. Our framework consists of a tokenizer that transforms continuous features into discrete tokens, and an LLM that is fine-tuned to predict motion tokens given the audio tokens. We show that by proper tokenization of input modalities and careful design of the LLM training strategies, our framework can generate realistic and diverse group dances while maintaining strong music correlation and dancer-wise consistency. Extensive experiments and evaluations demonstrate that our framework achieves state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。