比较四种生成模型在成人转儿童语音中的效果,提出频域校正提升相似性。
A comparative study of generative models for child voice conversion
- 对比扩散、流模型、变分自编码器与GAN在成人转儿童语音中的表现
- 模型输出语音虽自然但与目标儿童特征相似度不足,经频域校正后显著改善
- 适用于影视配音等需高保真儿童语音转换的场景
生成模型在成人到成人语音转换中表现优异,因其能高效建模无标签数据。然而其在生成儿童语音,特别是成人转儿童语音(adult-to-child VC)中的应用尚未被充分研究。本文对比了四种生成模型:扩散模型、基于流的模型、变分自编码器和生成对抗网络在成人转儿童语音任务中的表现。结果表明,尽管各模型生成的语音听起来自然可信,但与目标儿童说话人特征的相似度仍不理想。为此,本文提出一种高效的频域扭曲技术,可作用于模型输出,显著降低成人与儿童语音之间的差异。所有模型的输出均通过客观与主观评估进行验证,并使用一个专为儿童语音配音收集的独特语料库进行特定说话人配对比较。
原文摘要 · Abstract (English)
Generative models are a popular choice for adult-to-adult voice conversion (VC) because of their efficient way of modelling unlabelled data. To this point their usefulness in producing children speech and in particular adult to child VC has not been investigated. For adult to child VC, four generative models are compared: diffusion model, flow based model, variational autoencoders, and generative adversarial network. Results show that although converted speech outputs produce by those models appear plausible, they exhibit insufficient similarity with the target speaker characteristics. We introduce an efficient frequency warping technique that can be applied to the output of models, and which shows significant reduction of the mismatch between adult and child. The output of all the models are evaluated using both objective and subjective measures. In particular we compare specific speaker pairing using a unique corpus collected for dubbing of children speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。