用思维链提升动作生成与理解的可解释性与准确性
UniMo: Unified Motion Generation and Understanding with Chain of Thought
- 引入思维链推理,让语言模型更懂动作逻辑
- 在多个数据集上超越现有统一与专用模型表现
- 适合需要高可解释性的动作生成研究者使用
现有3D人体动作生成与理解方法普遍存在可解释性不足的问题,限制了两者间的协同提升。尽管当前基于大语言模型(LLM)的统一框架利用语言先验,却常面临语义对齐与任务一致性挑战。此外,LLM的逐词预测范式不适用于动作序列,导致累积预测误差。为此,我们提出UniMo,通过监督微调(SFT)将动作-语言信息与可解释的思维链(CoT)推理融入LLM。进一步采用分组相对策略优化(GRPO)作为后训练策略,对一组词元进行优化,以确保结构正确性和语义对齐,缓解动作词元预测中的累积误差。大量实验表明,UniMo在动作生成与理解任务上均显著优于现有统一及专用模型,达到当前最优性能。
原文摘要 · Abstract (English)
Existing 3D human motion generation and understanding methods often exhibit limited interpretability, restricting effective mutual enhancement between these inherently related tasks. While current unified frameworks based on large language models (LLMs) leverage linguistic priors, they frequently encounter challenges in semantic alignment and task coherence. Moreover, the next-token prediction paradigm in LLMs is ill-suited for motion sequences, causing cumulative prediction errors. To address these limitations, we propose UniMo, a novel framework that integrates motion-language information and interpretable chain of thought (CoT) reasoning into the LLM via supervised fine-tuning (SFT). We further introduce reinforcement learning with Group Relative Policy Optimization (GRPO) as a post-training strategy that optimizes over groups of tokens to enforce structural correctness and semantic alignment, mitigating cumulative errors in motion token prediction. Extensive experiments demonstrate that UniMo significantly outperforms existing unified and task-specific models, achieving state-of-the-art performance in both motion generation and understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。