arXiv:2409.13251cs.CV2024-09被引 5

用部分标注数据训练,生成包含面部和手部动作的全身动画。

T2M-X: Learning Expressive Text-to-Motion Generation from Partially Annotated Data

  • 分阶段训练身体、手部、面部三个VQ-VAE,保证高质量输出
  • 用多索引GPT模型生成动作并保持各部位协调一致
  • 在数据不完整时仍表现稳健,适合影视与VR场景

从文本提示生成类人动画可显著提升动画制作与AR/VR体验。然而,现有方法仅生成躯干动作,忽略面部表情和手部动作。这主要受限于缺乏全面的全身动作数据集。近期尝试构建此类数据集时,存在人工增强数据中不同身体部位运动不一致,或从RGB视频提取的数据质量较低的问题。本文提出T2M-X,一种两阶段方法,从部分标注数据中学习具表现力的文本到动作生成。T2M-X在各自高质量数据源上分别训练身体、手部、面部的三个VQ-VAE,以确保高质量动作输出,并采用带运动一致性损失的多索引生成式预训练变换器(GPT)模型实现动作生成与多部位协调。实验结果表明,该方法在定量与定性评估上均显著优于基线,展现出对数据集局限性的鲁棒性。

原文摘要 · Abstract (English)

The generation of humanoid animation from text prompts can profoundly impact animation production and AR/VR experiences. However, existing methods only generate body motion data, excluding facial expressions and hand movements. This limitation, primarily due to a lack of a comprehensive whole-body motion dataset, inhibits their readiness for production use. Recent attempts to create such a dataset have resulted in either motion inconsistency among different body parts in the artificially augmented data or lower quality in the data extracted from RGB videos. In this work, we propose T2M-X, a two-stage method that learns expressive text-to-motion generation from partially annotated data. T2M-X trains three separate Vector Quantized Variational AutoEncoders (VQ-VAEs) for body, hand, and face on respective high-quality data sources to ensure high-quality motion outputs, and a Multi-indexing Generative Pretrained Transformer (GPT) model with motion consistency loss for motion generation and coordination among different body parts. Our results show significant improvements over the baselines both quantitatively and qualitatively, demonstrating its robustness against the dataset limitations.

文本生成动作全身动画生成模型数据不完整

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。