仅用面部图像和无声肌电信号,就能生成多说话人的语音。
Speaking Without Sound: Multi-speaker Silent Speech Voicing with Facial Inputs Only
- 通过肌电信号提取语言内容,面部图像匹配说话人身份。
- 提出分离音高的内容嵌入,提升语言信息提取效果。
- 适合无声语音合成与个性化语音重建研究者。
本文提出一种新框架,无需任何可听输入即可生成多说话人语音。方法利用无声肌电(EMG)信号捕捉语言内容,同时使用面部图像匹配目标说话人的声线身份。特别地,我们设计了一种音高解耦的内容嵌入,有效提升了从肌电信号中提取语言信息的能力。大量实验表明,该方法可在无音频输入条件下生成多说话人语音,并验证了音高解耦策略的有效性。
原文摘要 · Abstract (English)
In this paper, we introduce a novel framework for generating multi-speaker speech without relying on any audible inputs. Our approach leverages silent electromyography (EMG) signals to capture linguistic content, while facial images are used to match with the vocal identity of the target speaker. Notably, we present a pitch-disentangled content embedding that enhances the extraction of linguistic content from EMG signals. Extensive analysis demonstrates that our method can generate multi-speaker speech without any audible inputs and confirms the effectiveness of the proposed pitch-disentanglement approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。