arXiv:2608.23279cs.CV2026-08中稿 · ICME 2026

提出分时空建模的扩散生成模型,实现人体动作细节可控生成

Spatiotemporally Decoupled Autoregressive Diffusion Model for Human Motion Generation

论文配图:Spatiotemporally Decoupled Autoregressive Diffusion Model for Human Motion Generation
图 1 · 摘自论文原文
  • 按关节独立编码,分离时空特征提升局部控制力
  • 在HumanML3D和KIT-ML上达到最优重建与生成效果
  • 适合需要精细动作编辑的应用场景

文本驱动的人体动作生成已取得显著进展,核心在于运动表示与生成架构。现有基于向量量化(VQ)的方法将动作压缩为离散标记,导致信息丢失,影响生成质量、多样性和泛化能力;而连续空间建模整体身体运动则限制了局部灵活性。扩散与自回归扩散模型虽表现优异,但对单个身体部位的细粒度控制仍不足。为此,我们提出统一的时空解耦框架DeMoDiff,重构表示与架构。设计关节级时空变分自编码器,分别编码每个身体关节,突破整体潜空间限制。结合时空掩码与注意力机制的自回归扩散生成器,兼顾生成能力与可编辑性。在HumanML3D和KIT-ML数据集上的大量实验表明,模型实现最先进的重建性能与出色的生成结果,且具备强大的时空编辑能力,充分验证其有效性。

原文摘要 · Abstract (English)

Text-driven human motion synthesis has made substantial development with two core modules of motion representation and generative architecture. For representation, Vector Quantization (VQ)-based methods compress motion data into discrete tokens while latent-based models operate directly in continuous space. However, both of these representations exhibit significant limitations. VQ-based methods suffer from inherent information loss, which compromises the quality, diversity, and generalization of generated motions, while continuous representation on holistic whole-body motion hinders part-level flexibility. For architecture, diffusion and autoregressive diffusion models have demonstrated their superiority, yet the fine-grained controllability over individual body parts is also limited. Thus, we propose a unified spatiotemporally decoupled framework named DeMoDiff, which jointly redesigns representation and architecture. To enhance representation extraction capabilities and offer greater part-level controllability, we present a spatial-temporal VAE that encodes each body joint rather than compressing the whole-body motion into a single latent space. Then, we incorporate spatial-temporal masking and attention mechanisms into an autoregressive diffusion generator, achieving both generative capability and controllable editability. Extensive experiments on the HumanML3D and KIT-ML datasets demonstrate that our model achieves state-of-the-art reconstruction performance and compelling motion generation results. Moreover, our framework demonstrates strong temporal and spatial editing capabilities, further validating its effectiveness. Our project page: https://rex0191.github.io/DeMoDiff/

动作生成扩散模型时空解耦自回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。