多尺度粗到精建模,实现快速精准的人体动作控制生成
Multi-scale Coarse-to-fine Modeling for Test-time Human Motion Control

- 分层离散化动作,从粗到精逐级预测完整动作序列
- 相比扩散模型快10倍,文本到动作生成FID降低48%,误差减少61%
- 适合需要实时交互和高精度控制的动作生成场景
我们提出MSCoT,一种用于测试时人体动作合成与控制的多尺度、粗到精模型。不同于依赖多次去噪或特定控制信号模块的现有方法,MSCoT将动作离散化为多尺度层级表示,并以粗到精方式在每个时间尺度上一次性预测整个标记序列。基于此范式,我们设计了一种高效的多尺度标记引导策略,克服离散采样难题,引导标记分布向控制目标收敛,实现快速灵活控制。为解决离散码本的局限性,引入轻量级标记重构器,在离散标记嵌入上添加连续残差,支持可微的测试时优化,确保与控制目标精确对齐。MSCoT在生成质量与控制一致性方面表现优异,推理速度较扩散模型提升10倍。在HumanML3D等主流基准上的实验表明,其在可控文本到动作生成任务中达到当前最优性能,运动质量提升48% FID,平均误差降低61%,推理速度提升10倍。
原文摘要 · Abstract (English)
We present MSCoT, a multi-scale, coarse-to-fine model for test-time human motion synthesis and control. Unlike recent approaches that rely on multiple iterative denoising/token-prediction steps, or modules tailored for specific control signals, MSCoT discretizes motion into a multi-scale hierarchical representation and predicts the entire token sequence at each temporal scale in a coarse-to-fine fashion. Building on this coarse-to-fine paradigm, we propose an efficient multi-scale token guidance strategy that overcomes the challenge of discrete sampling and steers the token distribution towards the control goals, allowing for fast and flexible control. To address the limitations of a discrete codebook, a lightweight token refiner further adds continuous residuals to the discrete token embeddings and allows differentiable test-time refinement optimization to ensure precise alignment with the control objectives. MSCoT is able to produce quality motions, consistent with the control constraints, while offering substantially faster sampling than diffusion-based approaches. Experiments on popular benchmarks demonstrate state-of-the-art controllable text-to-motion generation performance of MSCoT over existing baselines, with better motion quality (48% FID improvement), higher control accuracy (-61% avg error), and $10 \times$ faster inference speed on HumanML3D.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。