arXiv:2410.06787eess.AS2024-10中稿 · 2024 IEEE 20th Int…被引 1

改进FastPitch模型,实现罗马尼亚语自然语音合成与说话人适配

Efficient training strategies for natural sounding speech synthesis and speaker adaptation based on FastPitch

  • 基于FastPitch架构,优化训练策略以支持多说话人
  • 可合成已知及未知说话人语音,匿名身份也能生成清晰语音
  • 适用于需要去身份化或快速适配新说话人的场景

本文将FastPitch模型功能拓展至罗马尼亚语,将说话人数量从1人扩展至18人,实现了使用匿名身份进行语音合成,并能通过短参考样本复现未见过的说话人音色。研究测试了多种配置与训练策略,分析其优劣,最终确定一种新配置,可在已知和未知说话人条件下生成自然语音。匿名说话人模式可用于文本到语音合成,消除说话人特征而保留语义完整性。最后讨论了当前方法的局限性,为后续研究提供基础。

原文摘要 · Abstract (English)

This paper focuses on adapting the functionalities of the FastPitch model to the Romanian language; extending the set of speakers from one to eighteen; synthesising speech using an anonymous identity; and replicating the identities of new, unseen speakers. During this work, the effects of various configurations and training strategies were tested and discussed, along with their advantages and weaknesses. Finally, we settled on a new configuration, built on top of the FastPitch architecture, capable of producing natural speech synthesis, for both known (identities from the training dataset) and unknown (identities learnt through short reference samples) speakers. The anonymous speaker can be used for text-to-speech synthesis, if one wants to cancel out the identity information while keeping the semantic content whole and clear. At last, we discussed possible limitations of our work, which will form the basis for future investigations and advancements.

语音合成说话人适配FastPitch

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。