用扩散模型生成有创意又真实的舞蹈动作,让AI像人一样跳舞。
Generative human motion mimicking through feature extraction in denoising diffusion settings
- 基于单人动捕数据和高层特征,用扩散模型生成动作。
- 生成动作与真实人类数据特征分布接近,且多样性高。
- 适合想探索人机共舞的创作者或研究者。
大型语言模型的兴起推动了人机语音交互的发展,但缺乏具身性。舞蹈作为原始的人类表达形式,可弥补这一不足。为此,我们构建了一个基于动作捕捉(MoCap)数据的交互式模型,能部分模仿并创造性增强输入的动作序列。该模型首次利用单人运动数据和高层特征实现生成,无需依赖低级的人类间互动数据。它融合了两种扩散模型、动作补全与风格迁移思想,生成既时序连贯又响应参考动作的运动表示。通过量化评估生成样本与测试集特征分布的收敛性,验证了模型有效性。结果表明,生成动作在保持真实感的同时展现出多样化偏差,是迈向人机共创舞蹈的重要一步。
原文摘要 · Abstract (English)
Recent success with large language models has sparked a new wave of verbal human-AI interaction. While such models support users in a variety of creative tasks, they lack the embodied nature of human interaction. Dance, as a primal form of human expression, is predestined to complement this experience. To explore creative human-AI interaction exemplified by dance, we build an interactive model based on motion capture (MoCap) data. It generates an artificial other by partially mimicking and also "creatively" enhancing an incoming sequence of movement data. It is the first model, which leverages single-person motion data and high level features in order to do so and, thus, it does not rely on low level human-human interaction data. It combines ideas of two diffusion models, motion inpainting, and motion style transfer to generate movement representations that are both temporally coherent and responsive to a chosen movement reference. The success of the model is demonstrated by quantitatively assessing the convergence of the feature distribution of the generated samples and the test set which serves as simulating the human performer. We show that our generations are first steps to creative dancing with AI as they are both diverse showing various deviations from the human partner while appearing realistic.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。