arXiv:2604.08088cs.CV2026-04被引 1

用坐标约束提升动作生成质量,兼顾细节与语义一致性。

Coordinate-Based Dual-Constrained Autoregressive Motion Generation

  • 以运动坐标为输入,结合扩散模型思想增强生成细节
  • 提出双约束因果掩码,减少模式崩溃,提升生成多样性
  • 适用于动画、虚拟现实等需要高保真动作生成的场景

文本到动作生成近年来受到广泛关注,应用涵盖动画、虚拟现实、机器人及人机交互。扩散模型常因噪声预测导致误差累积,自回归模型则因动作离散化易出现模式崩溃。为此,本文提出一种灵活、高保真且语义忠实的文本到动作生成框架——基于坐标的双约束自回归动作生成(CDAMD)。该方法以运动坐标为输入,遵循自回归范式,并引入受扩散模型启发的多层感知机以提升生成动作的保真度。此外,设计了双约束因果掩码,将动作标记作为先验,与文本编码拼接,引导生成过程。针对坐标基动作合成研究较少的问题,本文建立了新的文本到动作生成与动作编辑基准。实验表明,所提方法在保真度与语义一致性上均达到当前最优水平。

原文摘要 · Abstract (English)

Text-to-motion generation has attracted increasing attention in the research community recently, with potential applications in animation, virtual reality, robotics, and human-computer interaction. Diffusion and autoregressive models are two popular and parallel research directions for text-to-motion generation. However, diffusion models often suffer from error amplification during noise prediction, while autoregressive models exhibit mode collapse due to motion discretization. To address these limitations, we propose a flexible, high-fidelity, and semantically faithful text-to-motion framework, named Coordinate-based Dual-constrained Autoregressive Motion Generation (CDAMD). With motion coordinates as input, CDAMD follows the autoregressive paradigm and leverages diffusion-inspired multi-layer perceptrons to enhance the fidelity of predicted motions. Furthermore, a Dual-Constrained Causal Mask is introduced to guide autoregressive generation, where motion tokens act as priors and are concatenated with textual encodings. Since there is limited work on coordinate-based motion synthesis, we establish new benchmarks for both text-to-motion generation and motion editing. Experimental results demonstrate that our approach achieves state-of-the-art performance in terms of both fidelity and semantic consistency on these benchmarks.

动作生成自回归模型坐标输入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。