分层生成动作,从粗到细提升文本驱动动画质量
Next-Scale Autoregressive Models for Text-to-Motion Generation

- 按粗粒度到细粒度分层生成动作,建立更符合时序结构的因果关系
- 在有限数据下仍保持高性能,支持零样本动作生成与编辑
- 训练高效且模型越大效果越好,适合大规模动作生成任务
自回归模型训练稳定高效,但标准的下一个词预测与文本条件动作生成所需的时序结构不匹配。我们提出MoScale,一种分层生成的动作自回归框架,从粗到细逐步生成运动序列。通过在最粗粒度提供全局语义,并逐级细化,建立更适配长时序结构的因果层次。为提升有限文本-动作数据下的鲁棒性,进一步引入跨尺度层次化精炼和同尺度时序精炼,实现选择性双向重预测。MoScale在文本到动作生成任务中达到当前最优性能,具备高训练效率,模型规模扩大时表现持续提升,并可零样本泛化至多样动作生成与编辑任务。
原文摘要 · Abstract (English)
Autoregressive (AR) models offer stable and efficient training, but standard next-token prediction is not well aligned with the temporal structure required for text-conditioned motion generation. We introduce MoScale, a next-scale AR framework that generates motion hierarchically from coarse to fine temporal resolutions. By providing global semantics at the coarsest scale and refining them progressively, MoScale establishes a causal hierarchy better suited for long-range motion structure. To improve robustness under limited text-motion data, we further incorporate cross-scale hierarchical refinement for improving per-scale initial predictions and in-scale temporal refinement for selective bidirectional re-prediction. MoScale achieves SOTA text-to-motion performance with high training efficiency, scales effectively with model size, and generalizes zero-shot to diverse motion generation and editing tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。