arXiv:2409.01591cs.CVcs.GR2024-09被引 1

用文本和音频联合生成自然连贯的全身动作序列

Dynamic Motion Synthesis: Masked Audio-Text Conditioned Spatio-Temporal Transformers

  • 通过向量量化自编码器离散化动作,结合掩码建模高效预测动作标记
  • 引入空间注意力与标记判别器,提升生成动作的一致性与自然度
  • 适合多模态动作生成、虚拟人驱动等场景

本研究提出一种新型动作生成框架,可同时基于文本和音频输入生成全身动作序列。通过向量量化变分自编码器(VQVAEs)对动作进行离散化,并采用双向掩码语言建模(MLM)策略实现高效的标记预测,显著提升处理效率与生成动作的连贯性。通过引入空间注意力机制与标记判别器,进一步保障生成动作在时空上的一致性和自然性。该框架突破了现有方法在多模态条件下的局限,拓展了动作生成的可能性,为多模态动作合成开辟新路径。

原文摘要 · Abstract (English)

Our research presents a novel motion generation framework designed to produce whole-body motion sequences conditioned on multiple modalities simultaneously, specifically text and audio inputs. Leveraging Vector Quantized Variational Autoencoders (VQVAEs) for motion discretization and a bidirectional Masked Language Modeling (MLM) strategy for efficient token prediction, our approach achieves improved processing efficiency and coherence in the generated motions. By integrating spatial attention mechanisms and a token critic we ensure consistency and naturalness in the generated motions. This framework expands the possibilities of motion generation, addressing the limitations of existing approaches and opening avenues for multimodal motion synthesis.

动作生成多模态变换器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。