arXiv:2512.10730cs.CV2025-12中稿 · ECCV被引 3

让动作生成、评估与优化交替进行,提升文本到动作的匹配度。

IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation

  • 通过文本-动作对话循环,交替生成、评估和优化动作。
  • 在多个基准上实现跨评测者性能提升,优于现有方法。
  • 适合需要高精度动作生成的研究者与开发者。

近期基于动作感知的大语言模型在联合学习动作理解与生成方面展现出巨大潜力。然而,这些模型通常将理解与生成任务分开处理,限制了两者间的互惠反馈。本文揭示,动作评估与优化可作为关键桥梁,促进从动作理解向生成的知识流动。为此,我们提出一种新的范式——交错推理动作生成(IRMoGen),通过迭代的文本-动作对话,紧密耦合动作生成、评估与优化。为实现该目标,我们构建了首个无缝交织生成、评估与优化的模型IRG-MotionLLM,采用新颖的三阶段训练方案逐步初始化并增强其原生能力。为支持开发,我们设计自动化数据引擎,从现有文本-动作数据集中合成交错推理标注。大量实验表明,IRMoGen训练带来了显著特性,且IRG-MotionLLM在跨基准与跨评测者场景中表现优异。代码与模型已公开于https://github.com/HumanMLLM/IRG-MotionLLM。

原文摘要 · Abstract (English)

Recent advances in motion-aware large language models have shown remarkable promise for jointly learning motion understanding and generation knowledge. However, these models typically treat understanding and generation separately, limiting the mutual benefits that could arise from interactive feedback between tasks. In this work, we reveal that motion assessment and refinement tasks can act as crucial bridges to enable knowledge flow from motion understanding to generation. Specifically, we propose Interleaved Reasoning for Motion Generation (IRMoGen), a novel paradigm that tightly couples motion generation with assessment and refinement through iterative text-motion dialogue. To realize this, we introduce IRG-MotionLLM, the first model that seamlessly interleaves motion generation, assessment, and refinement to improve the alignment between generated motion and goal text. IRG-MotionLLM is developed progressively with a novel three-stage training scheme, initializing and subsequently enhancing native IRMoGen capabilities. To facilitate this development, we construct an automated data engine to synthesize interleaved reasoning annotations from existing text-motion datasets. Extensive experiments demonstrate the properties brought by IRMoGen training, and the advanced cross-benchmark and cross-evaluator performance of IRG-MotionLLM. Code and models are available at https://github.com/HumanMLLM/IRG-MotionLLM.

动作生成大模型交互优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。