用空间音频生成逼真人体动作,首次构建专用数据集与模型
MOSPA: Human Motion Generation Driven by Spatial Audio
- 基于扩散模型融合空间音频特征生成动作
- 在自建数据集上实现最优生成质量
- 适合虚拟角色动画与交互系统研发者
让虚拟人物根据多样听觉刺激动态、真实地做出反应,是角色动画中的关键挑战,需结合感知建模与运动合成。尽管意义重大,该任务仍鲜有研究。以往工作多聚焦于语音、音乐等模态到动作的映射,却普遍忽略空间音频信号中编码的空间特征对动作的影响。为填补这一空白并实现高质量空间音频驱动的人体运动建模,我们首次构建了全面的时空音频驱动人体运动(SAM)数据集,包含多样化且高质量的空间音频与动作数据。为此,我们开发了一个简单而有效的基于扩散的生成框架MOSPA,通过高效融合机制忠实捕捉身体运动与空间音频之间的关系。训练完成后,MOSPA可依据不同空间音频输入生成多样且逼真的动作。我们对所提数据集进行了深入分析,并开展大量基准实验,结果表明本方法在该任务上达到当前最优性能。代码与模型已开源:https://github.com/xsy27/Mospa-Acoustic-driven-Motion-Generation
原文摘要 · Abstract (English)
Enabling virtual humans to dynamically and realistically respond to diverse auditory stimuli remains a key challenge in character animation, demanding the integration of perceptual modeling and motion synthesis. Despite its significance, this task remains largely unexplored. Most previous works have primarily focused on mapping modalities like speech, audio, and music to generate human motion. As of yet, these models typically overlook the impact of spatial features encoded in spatial audio signals on human motion. To bridge this gap and enable high-quality modeling of human movements in response to spatial audio, we introduce the first comprehensive Spatial Audio-Driven Human Motion (SAM) dataset, which contains diverse and high-quality spatial audio and motion data. For benchmarking, we develop a simple yet effective diffusion-based generative framework for human MOtion generation driven by SPatial Audio, termed MOSPA, which faithfully captures the relationship between body motion and spatial audio through an effective fusion mechanism. Once trained, MOSPA can generate diverse, realistic human motions conditioned on varying spatial audio inputs. We perform a thorough investigation of the proposed dataset and conduct extensive experiments for benchmarking, where our method achieves state-of-the-art performance on this task. Our code and model are publicly available at https://github.com/xsy27/Mospa-Acoustic-driven-Motion-Generation
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。