用语音生成实时声带运动影像,让说话过程可视化。
Speech2rtMRI: Speech-Guided Diffusion Model for Real-time MRI Video of the Vocal Tract during Speech
- 用语音驱动扩散模型生成声带动态视频
- 借助预训练语音模型提升生成质量
- 适合语言研究与虚拟角色动画开发
理解语音产生的视觉与运动特征,有助于改进第二语言学习系统设计,并推动视频游戏和动画中说话角色的创建。本文提出一种数据驱动方法,基于任意音频或语音输入,生成人声带在说话过程中进行磁共振成像(MRI)的视频。该方法利用具备先验知识的大规模预训练语音模型,通过语音到视频的扩散模型,将视觉域泛化至未见数据。实验表明,预训练语音表示显著提升了视觉生成效果。同时发现,孤立评估音素存在困难,而在真实词汇语境中则更易识别。当前结果的局限在于舌部运动不够平滑,以及舌头接触上颚时出现视频失真。
原文摘要 · Abstract (English)
Understanding speech production both visually and kinematically can inform second language learning system designs, as well as the creation of speaking characters in video games and animations. In this work, we introduce a data-driven method to visually represent articulator motion in Magnetic Resonance Imaging (MRI) videos of the human vocal tract during speech based on arbitrary audio or speech input. We leverage large pre-trained speech models, which are embedded with prior knowledge, to generalize the visual domain to unseen data using a speech-to-video diffusion model. Our findings demonstrate that the visual generation significantly benefits from the pre-trained speech representations. We also observed that evaluating phonemes in isolation is challenging but becomes more straightforward when assessed within the context of spoken words. Limitations of the current results include the presence of unsmooth tongue motion and video distortion when the tongue contacts the palate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。