用分步推理和强化学习提升文本生成动作的自然度与逻辑性。
Motion-R1: Enhancing Motion Generation with Decomposed Chain-of-Thought and RL Binding
- 分步推理框架捕捉语言与动作的时序因果关系
- 在HumanML3D上MM-Dist降低3.5%,多指标领先
- 适合需要高精度动作生成的交互系统研究者
文本到动作生成已成为人机交互的基础任务,可从自然语言描述中合成逼真人体动作。尽管大语言模型和强化学习的进步带来了高质量动作生成,但仍面临两大挑战:现有方法难以捕捉语言中的时序与因果复杂性,导致动作简化或不连贯;基于强化学习的方法通常过于复杂,限制了其在不同任务间的可扩展性与适应性。为此,我们提出Motion-R1框架,结合分解式思维链(Decomposed CoT)与强化学习,提升生成动作的质量与可解释性。具体地,引入自动化数据引擎生成高质量推理数据,使模型更好地理解动作的时序依赖与因果关系;提出RL Binding策略,将多模态文本-动作对齐纳入强化学习奖励函数,引导模型生成语义准确且动作真实的动作。在多个基准数据集上的实验表明,Motion-R1在HumanML3D上实现MM-Dist降低3.5%,并在KIT-ML和BABEL上显著提升R-Precision与FID,全面超越现有方法,展现出处理复杂动作生成任务的强大能力。
原文摘要 · Abstract (English)
Text-to-Motion generation has become a fundamental task in human-machine interaction, enabling the synthesis of realistic human motions from natural language descriptions. Although recent advances in large language models and reinforcement learning have contributed to high-quality motion generation, two major challenges remain. Existing approaches often fail to capture the temporal and causal complexities inherent in natural language, leading to oversimplified or incoherent motions. Additionally, RL-based methods are frequently overly complex, hindering their scalability and adaptability across various motion generation tasks. To address these challenges, we propose Motion-R1, a novel framework that combines decomposed Chain-of-Thought reasoning with reinforcement learning to enhance both the quality and interpretability of generated motions. Specifically, we introduce the Decomposed CoT Data Engine, which leverages an automated pipeline to synthesize high-quality reasoning data, allowing the model to better capture the temporal dependencies and causal relationships of human motion. We also propose RL Binding, a reinforcement learning strategy that incorporates multi-modal text-motion alignment into the RL reward function, guiding the model to produce motions that are both semantically accurate and motionally realistic. Extensive experiments across benchmark datasets demonstrate that Motion-R1 achieves state-of-the-art performance, with a 3.5% improvement in MM-Dist on HumanML3D and improvements in R-Precision and FID on KIT-ML and BABEL, surpassing existing methods across key metrics and highlighting its superior capability in handling complex motion generation tasks. Project page: https://motion-r1.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。