arXiv:2503.14040cs.GRcs.CV2025-03被引 2

无需离散编码,实现自然流畅的全身伴语手势生成。

MAG: Multi-Modal Aligned Autoregressive Co-Speech Gesture Generation without Vector Quantization

  • 用连续运动嵌入替代离散符号,避免信息丢失。
  • 结合音频文本融合与扩散建模,生成动作更真实多样。
  • 适合做高质量虚拟人、数字主播动作生成的研究者。

本文聚焦于全身伴语手势生成。现有方法通常采用自回归模型配合向量量化令牌生成手势,导致信息损失,影响动作真实性。为此,受真实人类运动连续性启发,提出MAG框架,无需离散分词即可实现高质量、多样化的伴语手势合成。首先引入运动-文本-音频对齐变分自编码器(MTA-VAE),利用预训练的WavCaps文本与音频嵌入增强语义与节奏对齐,提升动作真实感。其次提出多模态掩码自回归模型(MMAG),在不依赖向量量化的情况下,通过扩散模型实现连续运动嵌入的自回归建模,并引入混合粒度音频-文本融合模块作为扩散过程的条件,确保多模态一致性。在两个基准数据集上的大量实验表明,MAG在定量与定性指标上均达到当前最优表现,生成动作高度逼真且多样。代码将公开以促进后续研究。

原文摘要 · Abstract (English)

This work focuses on full-body co-speech gesture generation. Existing methods typically employ an autoregressive model accompanied by vector-quantized tokens for gesture generation, which results in information loss and compromises the realism of the generated gestures. To address this, inspired by the natural continuity of real-world human motion, we propose MAG, a novel multi-modal aligned framework for high-quality and diverse co-speech gesture synthesis without relying on discrete tokenization. Specifically, (1) we introduce a motion-text-audio-aligned variational autoencoder (MTA-VAE), which leverages pre-trained WavCaps' text and audio embeddings to enhance both semantic and rhythmic alignment with motion, ultimately producing more realistic gestures. (2) Building on this, we propose a multimodal masked autoregressive model (MMAG) that enables autoregressive modeling in continuous motion embeddings through diffusion without vector quantization. To further ensure multi-modal consistency, MMAG incorporates a hybrid granularity audio-text fusion block, which serves as conditioning for diffusion process. Extensive experiments on two benchmark datasets demonstrate that MAG achieves stateof-the-art performance both quantitatively and qualitatively, producing highly realistic and diverse co-speech gestures.The code will be released to facilitate future research.

手势生成扩散模型多模态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。