arXiv:2606.14721cs.GRcs.CV2026-06

将动作结构与细节分离生成,提升文本到动作的可控性与自然度。

DC-Motion: Decoupling Structure and Details via Discrete-Continuous Tokens for Human Motion Generation

论文配图:DC-Motion: Decoupling Structure and Details via Discrete-Continuous Tokens for Human Motion Generation
图 1 · 摘自论文原文
  • 用离散结构令牌+连续残差潜变量分解动作表示
  • 在HumanML3D上FID达14.8,优于扩散与离散基线
  • 适合需要精确时序控制的动作生成任务

文本到动作生成需同时建模全局动作结构与细微运动动态。现有方法多依赖连续扩散模型或向量量化离散表示:前者生成平滑动作但缺乏显式结构用于时间规划,后者提升可控性但将动作压缩进有限码本,损失精细动态。我们指出问题根源在于表征不匹配——动作语义(意图、阶段转换、时间布局)本质上是离散且可组合的,而关节轨迹与运动动态是连续且局部相关的。为此,提出DC-Motion,一种离散-连续因子化的人体动作生成框架。该框架将动作分解为捕捉全局动作布局的离散结构令牌和建模精细动态的连续残差潜变量。文本条件结构生成器通过迭代掩码建模预测离散令牌,扩散残差生成器则基于结构生成连续运动。在HumanML3D和KIT-ML上的实验表明,DC-Motion在FID和R-Precision上均表现优异,超越代表性扩散模型与离散令牌基线。

原文摘要 · Abstract (English)

Text-to-motion generation requires modeling both global action structure and fine-grained motion dynamics from natural language. Existing approaches typically rely on either continuous diffusion models or vector-quantized discrete representations. Diffusion models generate smooth motions but lack explicit compositional structure for temporal planning, while discrete token-based methods improve controllability but compress motion into finite codebooks, losing fine-grained dynamics. We argue that this limitation stems from a representation mismatch: action semantics such as intent, phase transitions, and temporal layout are inherently discrete and compositional, whereas joint trajectories and motion dynamics are continuous and locally correlated. To address this, we propose DC-Motion, a discrete-continuous factorized framework for human motion generation. DC-Motion decomposes motion into discrete structural tokens capturing global action layout and continuous residual latents modeling fine-grained dynamics. A text-conditioned structure generator predicts discrete tokens via iterative masked modeling, and a diffusion-based residual generator produces continuous motion conditioned on the structure. Experiments on HumanML3D and KIT-ML demonstrate that DC-Motion achieves strong performance in both FID and R-Precision, outperforming representative diffusion-based and discrete-token baselines.

动作生成离散连续文本驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。