用百万级动作数据训练,实现文本到动作的零样本生成。
Go to Zero: Towards Zero-shot Motion Generation with Million-scale Data
- 构建百万级高质量动作数据集MotionMillion,含200万条动作序列。
- 在零样本测试中表现优异,能生成复杂组合动作且泛化能力强。
- 适合做动作生成、智能动画与机器人控制的研究者参考。
基于文本描述生成多样且自然的人体动作序列是计算机视觉、图形学和机器人领域的重要挑战。尽管已有显著进展,当前方法在零样本泛化能力上仍受限于训练数据规模不足,且缺乏全面评估框架。本文提出通过构建大规模数据集与评测体系,推动文本到动作生成进入新阶段。我们设计高效标注流程,发布迄今为止最大的人体动作数据集MotionMillion,包含超过2000小时、200万条高质量动作序列;同时提出MotionMillion-Eval,作为最全面的零样本动作生成评估基准。采用可扩展架构,将模型规模扩展至70亿参数,并在该基准上验证性能。结果表明,模型在跨领域及复杂组合动作上具备强大泛化能力,标志着向真正零样本人体动作生成迈出关键一步。代码已开源。
原文摘要 · Abstract (English)
Generating diverse and natural human motion sequences based on textual descriptions constitutes a fundamental and challenging research area within the domains of computer vision, graphics, and robotics. Despite significant advancements in this field, current methodologies often face challenges regarding zero-shot generalization capabilities, largely attributable to the limited size of training datasets. Moreover, the lack of a comprehensive evaluation framework impedes the advancement of this task by failing to identify directions for improvement. In this work, we aim to push text-to-motion into a new era, that is, to achieve the generalization ability of zero-shot. To this end, firstly, we develop an efficient annotation pipeline and introduce MotionMillion-the largest human motion dataset to date, featuring over 2,000 hours and 2 million high-quality motion sequences. Additionally, we propose MotionMillion-Eval, the most comprehensive benchmark for evaluating zero-shot motion generation. Leveraging a scalable architecture, we scale our model to 7B parameters and validate its performance on MotionMillion-Eval. Our results demonstrate strong generalization to out-of-domain and complex compositional motions, marking a significant step toward zero-shot human motion generation. The code is available at https://github.com/VankouF/MotionMillion-Codes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。