arXiv:2512.19159cs.CV2025-12被引 2

用交错文本-动作指令统一生成人类动作,支持自由创作与编辑。

OmniMoGen: Unifying Human Motion Generation via Learning from Interleaved Text-Motion Instructions

  • 通过交错文本与动作指令,统一多种运动生成任务。
  • 在13.7万条数据上训练,新基准测试表现领先。
  • 支持组合编辑、自我反思生成,适合智能动画开发。

大型语言模型(LLMs)已将多种语言任务统一在一个框架中,但人类动作生成领域尚未实现类似统一。现有方法局限于孤立任务,难以支持自由形式和多目标生成。为此,我们提出OmniMoGen,一个通过交错文本-动作指令实现多样化动作生成的统一框架。基于简洁的RVQ-VAE与Transformer架构,OmniMoGen支持端到端指令驱动的动作生成。我们构建了X2Mo数据集,包含超过13.7万条交错文本-动作指令,并引入AnyContext基准用于评估交错动作生成能力。实验表明,OmniMoGen在文本到动作生成、动作编辑及AnyContext任务上均达到当前最优性能,展现出组合编辑、自我反思生成和知识引导生成等新兴能力。这些结果标志着迈向下一代智能动作生成的重要一步。项目页面:https://OmniMoGen.github.io/

原文摘要 · Abstract (English)

Large language models (LLMs) have unified diverse linguistic tasks within a single framework, yet such unification remains unexplored in human motion generation. Existing methods are confined to isolated tasks, limiting flexibility for free-form and omni-objective generation. To address this, we propose OmniMoGen, a unified framework that enables versatile motion generation through interleaved text-motion instructions. Built upon a concise RVQ-VAE and transformer architecture, OmniMoGen supports end-to-end instruction-driven motion generation. We construct X2Mo, a large-scale dataset of over 137K interleaved text-motion instructions, and introduce AnyContext, a benchmark for evaluating interleaved motion generation. Experiments show that OmniMoGen achieves state-of-the-art performance on text-to-motion, motion editing, and AnyContext, exhibiting emerging capabilities such as compositional editing, self-reflective generation, and knowledge-informed generation. These results mark a step toward the next intelligent motion generation. Project Page: https://OmniMoGen.github.io/.

动作生成多模态指令学习统一框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。