arXiv:2503.14919cs.CV2025-03ICCV被引 8

用多专家架构统一学习跨数据集人体动作,文本到动作生成更准更泛化。

GenM$^3$: Generative Pretrained Multi-path Motion Model for Text Conditional Human Motion Generation

  • 采用多专家VQ-VAE与多路径变换器,分别处理数据异质性和模态内差异。
  • 在220小时动作数据上训练,人类动作生成任务FID达0.035,刷新纪录。
  • 支持零样本迁移,适合需要强泛化能力的动作生成研究者使用。

扩大动作数据集对提升动作生成能力至关重要。然而,在大规模多源数据上训练会因动作内容差异带来数据异质性挑战。为此,我们提出生成式预训练多路径动作模型(GenM³),一个旨在学习统一动作表征的完整框架。GenM³包含两部分:1)多专家VQ-VAE(MEVQ-VAE),可适应不同数据集分布,学习统一的离散动作表征;2)多路径动作变换器(MMT),通过独立的模态专用路径增强模态内表征,每个路径包含密集激活的专家以应对模态内变化,并通过共享的文本-动作路径促进跨模态对齐。为支持大规模训练,我们整合并统一了11个高质量动作数据集(约220小时动作数据),并添加文本标注(近10,000条动作序列由大语言模型标注,300+条由人工专家标注)。在整合数据集上训练后,GenM³在HumanML3D基准上达到0.035的FID,显著优于现有方法;同时在IDEA400数据集上展现强大零样本泛化能力,证明其在多样化动作场景中的有效性与适应性。

原文摘要 · Abstract (English)

Scaling up motion datasets is crucial to enhance motion generation capabilities. However, training on large-scale multi-source datasets introduces data heterogeneity challenges due to variations in motion content. To address this, we propose Generative Pretrained Multi-path Motion Model (GenM\(^3\)), a comprehensive framework designed to learn unified motion representations. GenM\(^3\) comprises two components: 1) a Multi-Expert VQ-VAE (MEVQ-VAE) that adapts to different dataset distributions to learn a unified discrete motion representation, and 2) a Multi-path Motion Transformer (MMT) that improves intra-modal representations by using separate modality-specific pathways, each with densely activated experts to accommodate variations within that modality, and improves inter-modal alignment by the text-motion shared pathway. To enable large-scale training, we integrate and unify 11 high-quality motion datasets (approximately 220 hours of motion data) and augment it with textual annotations (nearly 10,000 motion sequences labeled by a large language model and 300+ by human experts). After training on our integrated dataset, GenM\(^3\) achieves a state-of-the-art FID of 0.035 on the HumanML3D benchmark, surpassing state-of-the-art methods by a large margin. It also demonstrates strong zero-shot generalization on IDEA400 dataset, highlighting its effectiveness and adaptability across diverse motion scenarios.

动作生成多模态扩散模型文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。