arXiv:2504.03639cs.CV2025-04CVPR被引 13

让文字生成动作时考虑人体体型差异,更真实自然。

Shape My Moves: Text-Driven Shape-Aware Synthesis of Human Motions

  • 用离散化动作令牌+连续体型信息,还原真实体型对应的运动
  • 在多个数据集上实现高保真、与文本匹配的体型感知动作生成
  • 适合需要真实人体差异的动作生成场景,如虚拟角色定制

现有文本到动作生成方法常忽略身体形态的影响,因学习统一标准体型较易。但这种同质化会扭曲不同体型与运动动态之间的自然关联。本文提出一种基于有限标量量化变分自编码器(FSQ-VAE)的方法,将动作离散化为令牌,并利用连续体型信息将其还原为详细连续动作。同时借助预训练语言模型,联合预测连续体型参数与动作令牌,实现文本对齐的动作合成与体型感知解码。我们在定量、定性及全面感知评估中验证了该方法在生成体型感知动作方面的有效性。

原文摘要 · Abstract (English)

We explore how body shapes influence human motion synthesis, an aspect often overlooked in existing text-to-motion generation methods due to the ease of learning a homogenized, canonical body shape. However, this homogenization can distort the natural correlations between different body shapes and their motion dynamics. Our method addresses this gap by generating body-shape-aware human motions from natural language prompts. We utilize a finite scalar quantization-based variational autoencoder (FSQ-VAE) to quantize motion into discrete tokens and then leverage continuous body shape information to de-quantize these tokens back into continuous, detailed motion. Additionally, we harness the capabilities of a pretrained language model to predict both continuous shape parameters and motion tokens, facilitating the synthesis of text-aligned motions and decoding them into shape-aware motions. We evaluate our method quantitatively and qualitatively, and also conduct a comprehensive perceptual study to demonstrate its efficacy in generating shape-aware motions.

动作生成体型感知文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。