用合成数据扩展动作表示,让生成模型学会更复杂多样的动作。
Beyond MoCap: Scaling Motion Tokenizers with Synthetic Human Motion for Generative Modeling

- 用大规模合成动作数据训练新型VQ-VAE分词器,扩大动作表征空间。
- 合成数据使动作词汇覆盖率和组合能力显著提升,生成效果更丰富。
- 适合做动作生成、文本到动作等任务的研究者,尤其关注泛化能力提升。
人体动作生成模型受限于动作捕捉数据集的多样性不足,这些数据主要包含常见重复动作,难以覆盖复杂、罕见及动态性强的动作,导致学习到的隐变量表示动作词汇有限,泛化能力差。本文提出一种利用大规模合成人体动作为基础扩展动作表征空间的框架,设计了生成多样且物理合理动作序列的数据流水线,并与重构的VQ-VAE分词器结合,使其适应更广的动作分布。不同于仅在窄分布上训练的传统分词器,本方法同时扩展训练分布与离散码本,使模型能捕捉更丰富的动作基元。实验表明,使用合成动作训练可显著提升动作词汇的覆盖范围与组合性,在文本到动作、动作续写等任务中均取得一致性能提升,且与MotionGPT等现有框架完全兼容。结果表明,主要瓶颈在于动作表示的支撑不足,而非模型架构本身。将合成动作与表征学习协同扩展,为实现更表达性强、可控性高、泛化能力强的人体动作合成提供了系统路径。
原文摘要 · Abstract (English)
Human motion generation models are fundamentally constrained by the limited diversity of motion capture datasets, which predominantly contain common, repetitive actions and fail to cover the long tail of complex human movements, resulting in a restricted motion vocabulary in learned latent representations and poor generalization to rare, compositional, and highly dynamic motions. In this work, we propose a framework for expanding the motion representation space by leveraging large-scale synthetic human motion, introducing a data generation pipeline that produces diverse, physically plausible motion sequences beyond the distribution of existing datasets and integrating it with a redesigned VQ-VAE tokenizer that adapts to this expanded motion space. Unlike conventional tokenizers trained on narrow data distributions, our approach jointly scales both the training distribution and the discrete codebook, enabling the model to capture a significantly richer set of motion primitives. We demonstrate that training with synthetic motion substantially improves the coverage and compositionality of the learned motion vocabulary, leading to consistent gains across motion generation tasks such as text-to-motion and motion continuation, while remaining fully compatible with existing frameworks including MotionGPT. Our results suggest that the primary bottleneck lies in the limited support of the learned motion representation, rather than model architecture alone. Scaling synthetic motion in tandem with representation learning offers a principled path toward more expressive, controllable, and generalizable human motion synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。