首个统一生成人与动物动作的文本驱动模型,提升跨物种动作真实性。
X-MoGen: Unified Motion Generation across Humans and Animals
- 分两阶段建模:先学通用姿态先验,再用掩码建模生成动作嵌入
- 在115种动物和人类上实现跨物种动作生成,未见物种也表现优异
- 自研UniMo4D数据集支持联合训练,骨架一致性模块保障动作合理
文本驱动的动作生成因在虚拟现实、动画和机器人领域的广泛应用而受到越来越多关注。现有方法通常分别建模人类和动物动作,而联合跨物种建模能提供统一表示并增强泛化能力。然而,物种间的形态差异仍是关键挑战,常导致动作不自然。为此,我们提出X-MoGen,首个覆盖人类与动物的统一跨物种文本驱动动作生成框架。该框架采用两阶段设计:第一阶段,通过条件图变分自编码器学习通用的T姿态先验,同时使用自编码器将动作编码到共享潜在空间,并由形态损失进行正则化;第二阶段,通过掩码动作建模生成基于文本描述的动作嵌入。训练中引入形态一致性模块以提升跨物种骨骼合理性。为支持统一建模,我们构建了大规模数据集UniMo4D,包含115个物种和119,000条动作序列,采用共享骨骼拓扑整合人类与动物动作,支持联合训练。在UniMo4D上的大量实验表明,X-MoGen在已见与未见物种上均优于现有最先进方法。
原文摘要 · Abstract (English)
Text-driven motion generation has attracted increasing attention due to its broad applications in virtual reality, animation, and robotics. While existing methods typically model human and animal motion separately, a joint cross-species approach offers key advantages, such as a unified representation and improved generalization. However, morphological differences across species remain a key challenge, often compromising motion plausibility. To address this, we propose X-MoGen, the first unified framework for cross-species text-driven motion generation covering both humans and animals. X-MoGen adopts a two-stage architecture. First, a conditional graph variational autoencoder learns canonical T-pose priors, while an autoencoder encodes motion into a shared latent space regularized by morphological loss. In the second stage, we perform masked motion modeling to generate motion embeddings conditioned on textual descriptions. During training, a morphological consistency module is employed to promote skeletal plausibility across species. To support unified modeling, we construct UniMo4D, a large-scale dataset of 115 species and 119k motion sequences, which integrates human and animal motions under a shared skeletal topology for joint training. Extensive experiments on UniMo4D demonstrate that X-MoGen outperforms state-of-the-art methods on both seen and unseen species.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。