arXiv:2411.17335cs.CV2024-11被引 7

一个统一框架,能生成理解人和多角色动作,并支持文本、音乐、语音间转换。

VersatileMotion: A Unified Framework for Motion Synthesis and Comprehension

  • 用新型动作分词器融合VQ-VAE与流匹配,结合自回归Transformer
  • 在九项任务中达七项领先性能,支持单/多智能体动作处理
  • 适合做动作生成与理解的通用基础模型,尤其跨模态应用

大型语言模型(LLMs)天生具备多任务学习能力:通过统一的下一个词预测范式,可自然应对多种下游任务。先前运动领域的工作已通过运动分词器与自回归Transformer适配LLMs,实现人类动作的生成与理解,但泛化能力有限,性能提升微弱。本文提出VersatileMotion,一种统一的多模态运动大模型,结合新型运动分词器(融合VQ-VAE与流匹配)和自回归Transformer主干,无缝支持至少九种不同的运动相关任务。VersatileMotion是首个在单一框架内处理单智能体与多智能体动作的方案,并实现动作、文本、音乐、语音间的跨模态转换,在七项任务上达到当前最优表现。MotionHub中的每个序列可包含自然语言描述、音乐或音频片段、语音转录及多智能体交互数据。为促进评估,我们定义并发布了覆盖九个核心任务的基准划分。大量实验表明,VersatileMotion在性能、泛化性和未来运动理解与生成方面的潜力均表现出色。

原文摘要 · Abstract (English)

Large language models (LLMs) are, by design, inherently capable of multi-task learning: through a unified next-token prediction paradigm, they can naturally address a wide variety of downstream tasks. Prior work in the motion domain has demonstrated some generality by adapting LLMs via a Motion Tokenizer coupled with an autoregressive Transformer to generate and understand human motion. However, this generality remains limited in scope and yields only modest performance gains. We introduce VersatileMotion, a unified multimodal motion LLM that combines a novel motion tokenizer, integrating VQ-VAE with flow matching, and an autoregressive transformer backbone to seamlessly support at least nine distinct motion-related tasks. VersatileMotion is the first method to handle single-agent and multi-agent motions in a single framework and enable cross-modal conversion between motion, text, music, and speech, achieving state-of-the-art performance on seven of these tasks. Each sequence in MotionHub may include one or more of the following annotations: natural-language captions, music or audio clips, speech transcripts, and multi-agent interaction data. To facilitate evaluation, we define and release benchmark splits covering nine core tasks. Extensive experiments demonstrate the superior performance, versatility, and potential of VersatileMotion as a foundational model for future understanding and generation of motion.

动作生成多模态统一框架大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。