千亿参数模型实现文本生成3D动作,精度远超开源基准。
HY-Motion 1.0: Scaling Flow Matching Models for Text-To-Motion Generation
- 基于DiT的流匹配模型,全阶段训练提升指令对齐
- 覆盖200+动作类别,生成质量显著优于现有开源模型
- 适合动作生成、游戏动画等需要高精度文本控制的场景
我们提出HY-Motion 1.0,一系列达到顶尖水平的大规模动作生成模型,可从文本描述生成3D人体动作。该模型是首个将基于扩散Transformer(DiT)的流匹配模型成功扩展至十亿参数量级的动作生成系统,在指令跟随能力上显著超越当前开源基准。其独特之处在于引入全面的全流程训练范式:在超过3000小时的动作数据上进行大规模预训练,在400小时精心筛选的数据上进行高质量微调,并结合人类反馈与奖励模型进行强化学习,确保文本指令与动作输出的高度一致。该框架依托严格的数据处理流程,实现精细的动作清洗与标注。最终模型覆盖6大类共200余种动作类别,具备最广范围的覆盖能力。我们已将HY-Motion 1.0开源,以推动后续研究并加速3D人体动作生成技术向商业化成熟迈进。
原文摘要 · Abstract (English)
We present HY-Motion 1.0, a series of state-of-the-art, large-scale, motion generation models capable of generating 3D human motions from textual descriptions. HY-Motion 1.0 represents the first successful attempt to scale up Diffusion Transformer (DiT)-based flow matching models to the billion-parameter scale within the motion generation domain, delivering instruction-following capabilities that significantly outperform current open-source benchmarks. Uniquely, we introduce a comprehensive, full-stage training paradigm -- including large-scale pretraining on over 3,000 hours of motion data, high-quality fine-tuning on 400 hours of curated data, and reinforcement learning from both human feedback and reward models -- to ensure precise alignment with the text instruction and high motion quality. This framework is supported by our meticulous data processing pipeline, which performs rigorous motion cleaning and captioning. Consequently, our model achieves the most extensive coverage, spanning over 200 motion categories across 6 major classes. We release HY-Motion 1.0 to the open-source community to foster future research and accelerate the transition of 3D human motion generation models towards commercial maturity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。