用分尺度自回归生成动作,细节更准,还能直接文本编辑。
ScaleMoGen: Autoregressive Next-Scale Prediction for Human Motion Generation

- 分层次预测动作片段,从粗到细逐步生成
- 在HumanML3D上FID达0.030,优于前人模型
- 无需训练即可通过文本修改动作,适合交互应用
我们提出ScaleMoGen,一种基于文本驱动的人体动作生成的分尺度自回归框架。与传统自回归方法依赖标准的下一个标记预测不同,ScaleMoGen将动作生成视为从粗到细的过程。我们将3D动作在多个逐渐细化的骨骼-时间尺度上量化为组合式离散标记,通过自回归预测下一尺度的标记图来生成动作。为保持结构完整性,我们的动作标记器和量化器被显式设计为在每个尺度上严格保留骨骼层级关系。此外,采用位级量化与预测,高效扩展标记器词汇量,以保留动作细节并稳定优化过程。大量实验表明,ScaleMoGen达到领先性能,在HumanML3D上取得0.030的FID(前人0.045),在SnapMoGen数据集上获得0.693的CLIP Score(前人0.685)。此外,我们证明了该骨骼-时间多尺度表示可自然支持无需训练的文本引导动作编辑。
原文摘要 · Abstract (English)
We present ScaleMoGen, a scale-wise autoregressive framework for text-driven human motion generation. Unlike conventional autoregressive approaches that rely on standard next-token prediction, ScaleMoGen frames motion generation as a coarse-to-fine process. We quantize 3D motions into compositional discrete tokens across multiple skeletal-emporal scales of increasing granularity, learning to generate motion by autoregressively predicting next-scale token maps. To maintain structural integrity, our motion tokenizers and quantizers are explicitly designed so that discrete tokens at every scale strictly preserve the skeletal hierarchy. Additionally, we employ bitwise quantization and prediction, which efficiently scale up the tokenizer vocabulary to preserve motion details and stabilize optimization. Extensive experiments demonstrate that ScaleMoGen achieves state-of-the-art performance, establishing an FID of 0.030 (vs. 0.045 for MoMask) on HumanML3D and a CLIP Score of 0.693 (vs. 0.685 for MoMask++) on the SnapMoGen dataset. Furthermore, we demonstrate that our skeletal-temporal multi-scale representation naturally facilitates training-free, text-guided motion editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。