用多个智能体生成复杂人体动作,支持细粒度控制和文本编辑。
CoMA: Compositional Human Motion Generation with Multi-modal Agents
- 多智能体协作,结合大模型与掩码Transformer实现动作生成。
- 在HumanML3D上表现优于现有方法,长序列动作生成更自然。
- 适合需要精细动作编辑与复杂场景生成的研究者或创作者。
近年来,3D人体动作生成取得了显著进展。尽管顶尖方法性能大幅提升,但在训练数据中未见的复杂、细节丰富的动作仍难以生成,主要受限于动作数据集稀缺及新数据生成成本高昂。为此,我们提出CoMA,一种基于智能体的复杂人体动作生成、编辑与理解框架。CoMA利用多个由大语言模型和视觉模型驱动的协同智能体,结合基于掩码Transformer的动作生成器,该生成器采用身体部位专用编码器与代码本,实现细粒度控制。该框架支持以详细指令生成短/长动作序列,实现文本引导的动作编辑,并具备自我纠错能力以提升质量。在HumanML3D数据集上的评估表明,其性能媲美当前最优方法。此外,我们构建了一组上下文丰富、组合性强且较长的文本提示,用户研究表明,该方法显著优于现有方案。
原文摘要 · Abstract (English)
3D human motion generation has seen substantial advancement in recent years. While state-of-the-art approaches have improved performance significantly, they still struggle with complex and detailed motions unseen in training data, largely due to the scarcity of motion datasets and the prohibitive cost of generating new training examples. To address these challenges, we introduce CoMA, an agent-based solution for complex human motion generation, editing, and comprehension. CoMA leverages multiple collaborative agents powered by large language and vision models, alongside a mask transformer-based motion generator featuring body part-specific encoders and codebooks for fine-grained control. Our framework enables generation of both short and long motion sequences with detailed instructions, text-guided motion editing, and self-correction for improved quality. Evaluations on the HumanML3D dataset demonstrate competitive performance against state-of-the-art methods. Additionally, we create a set of context-rich, compositional, and long text prompts, where user studies show our method significantly outperforms existing approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。