用强化学习提升动作理解与生成的推理能力,支持逐步思考与反思。
MoRL: Reinforced Reasoning for Unified Motion Understanding and Generation
- 结合监督微调与可验证奖励的强化学习,统一建模动作理解与生成。
- 在HumanML3D和KIT-ML上超越现有基线,显著提升逻辑连贯性与物理合理性。
- 提出链式动作推理(CoM),适合需要高精度动作规划的机器人应用。
人体动作理解与生成对视觉和机器人技术至关重要,但现有方法在推理能力和测试时规划方面仍受限。本文提出MoRL,一种通过监督微调与可验证奖励的强化学习训练的统一多模态动作模型。任务特定奖励设计融合语义对齐与推理连贯性(用于理解),以及物理合理性与文本-动作一致性(用于生成),显著提升逻辑推理与感知真实性。为增强推理能力,引入测试时推理方法链式动作(CoM),实现分步规划与反思。同时构建两个大规模思维链数据集:MoUnd-CoT-140K(理解)与MoGen-CoT-140K(生成),使动作序列与推理轨迹、动作描述对齐。在HumanML3D与KIT-ML数据集上的实验表明,MoRL在多项指标上优于当前最优基线。代码与网站已开源。
原文摘要 · Abstract (English)
Human motion understanding and generation are crucial for vision and robotics but remain limited in reasoning capability and test-time planning. We propose MoRL, a unified multimodal motion model trained with supervised fine-tuning and reinforcement learning with verifiable rewards. Our task-specific reward design combines semantic alignment and reasoning coherence for understanding with physical plausibility and text-motion consistency for generation, improving both logical reasoning and perceptual realism. To further enhance inference, we introduce Chain-of-Motion (CoM), a test-time reasoning method that enables step-by-step planning and reflection. We also construct two large-scale CoT datasets, MoUnd-CoT-140K and MoGen-CoT-140K, to align motion sequences with reasoning traces and action descriptions. Experiments on HumanML3D and KIT-ML show that MoRL achieves significant gains over state-of-the-art baselines. Code: https://github.com/AIGeeksGroup/MoRL. Website: https://aigeeksgroup.github.io/MoRL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。