通过模型编辑在神经网络参数空间生成多样语音,实现无参考音色合成。
Eigenvoice Synthesis based on Model Editing for Speaker Generation
- 在DNN参数空间定义说话人空间,直接采样生成新语音
- 实验验证可生成多样化说话人语音,发现性别主导轴
- 适合语音合成、说话人控制研究者使用
说话人生成任务旨在不依赖参考语音的情况下生成未见说话人的声音。该任务的关键在于定义能代表多样说话人的说话人空间,以确定生成语音的特征。然而,如何有效定义这一空间仍不明确。传统的基于参数化合成框架(如基于HMM的方法)中,Eigenvoice合成是一种有前景的方案,它利用预存的说话人特征定义低维说话人空间。本文提出一种基于模型编辑的新型DNN-Eigenvoice合成方法。与以往方法不同,本方法在DNN模型参数空间中定义说话人空间。通过直接在此空间中采样新的DNN模型参数,即可生成多样化的说话人语音。实验结果表明,该方法具备生成多样化说话人语音的能力。此外,我们发现了所构建说话人空间中的性别主导轴,表明该方法具有控制说话人属性的潜力。
原文摘要 · Abstract (English)
Speaker generation task aims to create unseen speaker voice without reference speech. The key to the task is defining a speaker space that represents diverse speakers to determine the generated speaker trait. However, the effective way to define this speaker space remains unclear. Eigenvoice synthesis is one of the promising approaches in the traditional parametric synthesis framework, such as HMM-based methods, which define a low-dimensional speaker space using pre-stored speaker features. This study proposes a novel DNN-based eigenvoice synthesis method via model editing. Unlike prior methods, our method defines a speaker space in the DNN model parameter space. By directly sampling new DNN model parameters in this space, we can create diverse speaker voices. Experimental results showed the capability of our method to generate diverse speakers' speech. Moreover, we discovered a gender-dominant axis in the created speaker space, indicating the potential to control speaker attributes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。