让动作生成更懂复杂度,动态帧重点处理
Not All Frames Are Equal: Complexity-Aware Masked Motion Generation via Motion Spectral Descriptors
- 用运动频谱描述符衡量局部动态复杂度
- 复杂动作生成误差降低,整体FID提升12.7%
- 无需训练,适合需要精准动作生成的场景
掩码生成模型在文本到动作合成中表现强劲,但对动作帧的处理过于均匀,未能匹配动作本身动态复杂度随时间剧烈变化的特点。我们发现当前方法在动态复杂的动作上性能显著下降,且帧级生成误差与运动动力学高度相关。为此,提出运动频谱描述符(MSD),一种基于运动速度短时谱计算的、无参数且可解释的局部动态复杂度度量。利用MSD实现复杂度感知的掩码生成:训练时指导内容聚焦的掩码策略,推理时提供频谱相似性先验增强自注意力,并可在迭代解码中调节词元采样。基于掩码生成器构建的DynMask方法,在HumanML3D和KIT-ML数据集上,复杂动作生成质量明显提升,整体FID分别改善12.7%和8.3%。结果表明,尊重局部运动复杂度是掩码动作生成的重要设计原则。
原文摘要 · Abstract (English)
Masked generative models have become a strong paradigm for text-to-motion synthesis, but they still treat motion frames too uniformly during masking, attention, and decoding. This is a poor match for motion, where local dynamic complexity varies sharply over time. We show that current masked motion generators degrade disproportionately on dynamically complex motions, and that frame-wise generation error is strongly correlated with motion dynamics. Motivated by this mismatch, we introduce the Motion Spectral Descriptor (MSD), a simple and parameter-free measure of local dynamic complexity computed from the short-time spectrum of motion velocity. Unlike learned difficulty predictors, MSD is deterministic, interpretable, and derived directly from the motion signal itself. We use MSD to make masked motion generation complexity-aware. In particular, MSD guides content-focused masking during training, provides a spectral similarity prior for self-attention, and can additionally modulate token-level sampling during iterative decoding. Built on top of masked motion generators, our method, DynMask, improves motion generation most clearly on dynamically complex motions while also yielding stronger overall FID on HumanML3D and KIT-ML. These results suggest that respecting local motion complexity is a useful design principle for masked motion generation. Project page: https://xiangyue-zhang.github.io/DynMask
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。