让虚拟人根据语音生成自然全身动作,支持不同节奏的肢体表达。
M3G: Multi-Granular Gesture Generator for Audio-Driven Full-Body Human Motion Synthesis
- 用可变粒度的编码器将动作分层表示,适应不同长度的动作模式。
- 在多个数据集上生成动作更自然,主观评价优于现有方法。
- 适合虚拟主播、游戏角色等需要语音驱动全身动作的场景。
从音频生成包含面部、身体、双手和整体运动的完整人体姿态,是虚拟形象构建中的重要但具挑战性的任务。以往方法对动作进行逐帧标记并基于输入音频预测每帧标记,但未考虑不同动作模式所需的帧数差异(即粒度不一)。现有方法因固定粒度的标记机制,无法有效建模多样化的动作模式。为此,我们提出多粒度动作生成框架 M3G。M3G 首创多粒度向量量化自编码器(MGVQ-VAE),实现不同时间粒度下的动作模式标记与序列重建;设计多粒度标记预测器,从音频中提取多尺度信息并预测对应的动作标记;最终通过 MGVQ-VAE 重构完整人体动作。客观与主观实验均表明,M3G 在生成自然且富有表现力的全身体态方面优于当前最优方法。
原文摘要 · Abstract (English)
Generating full-body human gestures encompassing face, body, hands, and global movements from audio is a valuable yet challenging task in virtual avatar creation. Previous systems focused on tokenizing the human gestures framewisely and predicting the tokens of each frame from the input audio. However, one observation is that the number of frames required for a complete expressive human gesture, defined as granularity, varies among different human gesture patterns. Existing systems fail to model these gesture patterns due to the fixed granularity of their gesture tokens. To solve this problem, we propose a novel framework named Multi-Granular Gesture Generator (M3G) for audio-driven holistic gesture generation. In M3G, we propose a novel Multi-Granular VQ-VAE (MGVQ-VAE) to tokenize motion patterns and reconstruct motion sequences from different temporal granularities. Subsequently, we proposed a multi-granular token predictor that extracts multi-granular information from audio and predicts the corresponding motion tokens. Then M3G reconstructs the human gestures from the predicted tokens using the MGVQ-VAE. Both objective and subjective experiments demonstrate that our proposed M3G framework outperforms the state-of-the-art methods in terms of generating natural and expressive full-body human gestures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。