用少量数据生成多样人类动作,还能同步语音生成手势。
Multi-Resolution Generative Modeling of Human Motion from Limited Data
- 多尺度架构结合骨骼卷积,分层生成不同帧率的动作
- 在有限数据下仍能生成多样化动作,全局与局部多样性均优
- 可直接合成SMPL姿态参数,适合动作生成与语音驱动手势应用
我们提出一种生成模型,从有限训练序列中学习合成人类动作。该框架支持跨多时间分辨率的条件生成与动作融合。通过整合骨骼卷积层与多尺度结构,模型有效捕捉人体运动模式。模型包含一系列生成网络与对抗网络,以及针对特定帧率设计的嵌入模块,实现对动作内容与细节的精确控制。值得注意的是,该方法还可扩展至共言语手势合成,即使在配对数据稀缺的情况下,也能从语音输入生成同步手势。通过直接合成SMPL姿态参数,该方法避免了测试时对人形网格的调整。实验表明,模型在覆盖训练样本方面表现优异,且生成动作具有高度多样性,本地与全局多样性指标均表现良好。
原文摘要 · Abstract (English)
We present a generative model that learns to synthesize human motion from limited training sequences. Our framework provides conditional generation and blending across multiple temporal resolutions. The model adeptly captures human motion patterns by integrating skeletal convolution layers and a multi-scale architecture. Our model contains a set of generative and adversarial networks, along with embedding modules, each tailored for generating motions at specific frame rates while exerting control over their content and details. Notably, our approach also extends to the synthesis of co-speech gestures, demonstrating its ability to generate synchronized gestures from speech inputs, even with limited paired data. Through direct synthesis of SMPL pose parameters, our approach avoids test-time adjustments to fit human body meshes. Experimental results showcase our model's ability to achieve extensive coverage of training examples, while generating diverse motions, as indicated by local and global diversity metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。