arXiv:2508.08991cs.CV2025-08被引 3

提出多尺度量化方法,让动作生成更灵活、能组合不同动作片段。

Spatial-Temporal Multi-Scale Quantization for Flexible Motion Generation

  • 用多尺度离散令牌表示动作,兼顾空间和时间粒度。
  • 无需重训练即可组合不同动作,支持编辑与控制。
  • 在多个基准上优于现有方法,适合复杂动作生成任务。

尽管人类动作生成取得显著进展,当前的运动表示通常以离散帧序列形式呈现,仍存在两大关键局限:(i) 缺乏多尺度视角,限制了对复杂运动模式的建模能力;(ii) 组合灵活性不足,影响模型在多样生成任务中的泛化性能。为此,我们提出 MSQ——一种新型量化方法,将运动序列压缩为跨空间与时间维度的多尺度离散令牌。MSQ 使用不同编码器捕捉身体部位在不同空间粒度下的特征,并在时间上插值后,在多个尺度上进行量化。基于此表示,我们构建了一个生成掩码建模模型,有效支持动作编辑、动作控制与条件生成。定量与定性分析表明,该量化方法可实现运动令牌的无缝组合,无需特殊设计或重新训练。大量实验验证了本方法在多个基准上的优越性。

原文摘要 · Abstract (English)

Despite significant advancements in human motion generation, current motion representations, typically formulated as discrete frame sequences, still face two critical limitations: (i) they fail to capture motion from a multi-scale perspective, limiting the capability in complex patterns modeling; (ii) they lack compositional flexibility, which is crucial for model's generalization in diverse generation tasks. To address these challenges, we introduce MSQ, a novel quantization method that compresses the motion sequence into multi-scale discrete tokens across spatial and temporal dimensions. MSQ employs distinct encoders to capture body parts at varying spatial granularities and temporally interpolates the encoded features into multiple scales before quantizing them into discrete tokens. Building on this representation, we establish a generative mask modeling model to effectively support motion editing, motion control, and conditional motion generation. Through quantitative and qualitative analysis, we show that our quantization method enables the seamless composition of motion tokens without requiring specialized design or re-training. Furthermore, extensive evaluations demonstrate that our approach outperforms existing baseline methods on various benchmarks.

动作生成多尺度量化灵活生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。