arXiv:2604.17090cs.CV2026-04

用骨骼坐标统一动作识别与文本生成,提升运动建模效果。

Marrying Text-to-Motion Generation with Skeleton-Based Action Recognition

论文配图:Marrying Text-to-Motion Generation with Skeleton-Based Action Recognition
图 1 · 摘自论文原文
  • 基于骨骼坐标实现从粗到细的运动扩散生成
  • 多模态识别器提供语义梯度指导,提升生成质量
  • 支持动作识别、文本生成等四类任务,通用性强

人体动作识别与运动生成是以人为中心计算机视觉中的两个活跃研究方向,均致力于使运动与文本语义对齐。然而,现有工作大多分别研究这两个问题,未揭示其内在联系——运动生成需依赖语义理解。本文通过骨骼坐标统一动作理解与生成,提出基于坐标的自回归运动扩散模型(CoAMD)。CoAMD采用从粗到细的生成策略,其核心为多模态动作识别器(MAR),可提供梯度级语义引导。此外,我们建立严格基准,在绝对坐标下评估基线模型。所提方法可应用于四类重要任务:基于骨骼的动作识别、文本到运动生成、文本-运动检索及运动编辑。在13个跨任务基准上的大量实验表明,该方法达到当前最优性能,验证了其在人体运动建模中的有效性与通用性。代码已开源。

原文摘要 · Abstract (English)

Human action recognition and motion generation are two active research problems in human-centric computer vision, both aiming to align motion with textual semantics. However, most existing works study these two problems separately, without uncovering the links between them, namely that motion generation requires semantic comprehension. This work investigates unified action recognition and motion generation by leveraging skeleton coordinates for both motion understanding and generation. We propose Coordinates-based Autoregressive Motion Diffusion (CoAMD), which synthesizes motion in a coarse-to-fine manner. As a core component of CoAMD, we design a Multi-modal Action Recognizer (MAR) that provides gradient-based semantic guidance for motion generation. Furthermore, we establish a rigorous benchmark by evaluating baselines on absolute coordinates. Our model can be applied to four important tasks, including skeleton-based action recognition, text-to-motion generation, text-motion retrieval, and motion editing. Extensive experiments on 13 benchmarks across these tasks demonstrate that our approach achieves state-of-the-art performance, highlighting its effectiveness and versatility for human motion modeling. Code is available at https://github.com/jidongkuang/CoAMD.

动作生成骨骼建模多模态扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。