用语音转换技术实现逼真变声,可模仿他人声音
Speech to Speech Synthesis for Voice Impersonation
- 融合语音识别与合成技术,实现跨风格语音转换
- 生成音频自然度高,优于对比的生成对抗模型
- 适合语音仿冒、角色配音等需要变声的应用
尽管语音识别和语音合成模型已取得显著进展,但语音到语音处理模型仍研究不足。本文提出语音到语音合成网络(STSSN),基于当前最先进系统,融合两大学科,实现高效的语音风格迁移以达成语音仿冒目的。实验表明,该模型虽存在容量限制,但仍能生成高度逼真的音频样本。通过与同类任务的生成对抗模型对比,本模型在生成效果上更具说服力。
原文摘要 · Abstract (English)
Numerous models have shown great success in the fields of speech recognition as well as speech synthesis, but models for speech to speech processing have not been heavily explored. We propose Speech to Speech Synthesis Network (STSSN), a model based on current state of the art systems that fuses the two disciplines in order to perform effective speech to speech style transfer for the purpose of voice impersonation. We show that our proposed model is quite powerful, and succeeds in generating realistic audio samples despite a number of drawbacks in its capacity. We benchmark our proposed model by comparing it with a generative adversarial model which accomplishes a similar task, and show that ours produces more convincing results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。