arXiv:2608.24334cs.CVcs.CL2026-08

让动作生成更懂语义,用分层编码提升文本到动作的生成质量

SeMoCo: A Semantic-First Motion Codec for Motion Language Modeling

论文配图:SeMoCo: A Semantic-First Motion Codec for Motion Language Modeling
图 1 · 摘自论文原文
  • 动作令牌分语义与运动细节两部分,优先保证语义准确
  • 在重建精度上优于现有编码器,生成动作更自然
  • 适合做动作生成、动画设计的开发者和研究者

离散动作表示已显著推动自回归文本到动作生成的发展。然而,大多数动作分词器以重建为目标优化,未根据语义角色分配容量。因此,动作级语义和精细运动细节必须通过同一重建驱动的层次结构进行编码。本文提出SeMoCo,一种以语义为先的动作编码器,以及用于语言条件动作生成的双轴生成器。每个动作令牌包含一个语义令牌和一组残差运动令牌。生成器建模时间维度上的语义进展,并自回归地细化残差项。我们还构建了Ω-MotionVerse,一个大规模、多源的人类动作数据集,统一采用SOMA表示。在各项对比中,SeMoCo在重建精度上优于现有编码器,其动作令牌在下游生成任务中表现优异。

原文摘要 · Abstract (English)

Discrete motion representations have substantially advanced autoregressive text-to-motion generation. However, most motion tokenizers are optimized for reconstruction and do not explicitly allocate capacity according to semantic role. Action-level meaning and fine-grained kinematic detail must therefore be encoded through the same reconstruction-driven hierarchy. We introduce SeMoCo, a semantic-first motion codec, together with a dual-axis motion generator for language-conditioned motion generation. Each motion token contains one semantic token and a residual sequence of kinematic tokens. The generator models semantic progression across time and autoregressively refines the residual entries. We also construct $Ω$-MotionVerse, a large-scale, multi-source human-motion dataset unified under the SOMA representation. Across the reported comparisons, SeMoCo achieves the best reconstruction accuracy among the compared codecs, while strong text-to-motion results demonstrate the effectiveness of its motion tokens for downstream generation.

动作生成语义编码扩散模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。