用离散语音单元实现多语调一键转换,保留说话人特征
Accent conversion using discrete units with parallel data synthesized from controllable accented TTS
- 用自监督语音聚类生成离散单元作为中间表示
- 转换后发音更自然,能准确保留原说话人特征
- 无需大量非母语数据,适合多语种语音迁移
音调转换(AC)的目标是改变语音语调同时保持内容和说话人身份。以往方法在推理时需参考语句,难以保持说话人特征,或仅支持一对一训练,无法扩展到多种非母语口音。本文提出一种可转换多种口音为母语口音的高效模型。该方法利用自监督语音表示聚类生成离散单元,作为音调转换的中间目标;通过多说话人文本转语音合成,将离散单元还原为自然母语发音,并保留原始说话人特征。此外,我们设计了一种高效的无监督数据增强方法,显著减少对非母语资源的需求。实验表明,该系统有效提升非母语者发音流畅度,听起来像母语者,且说话人特征保留良好。
原文摘要 · Abstract (English)
The goal of accent conversion (AC) is to convert speech accents while preserving content and speaker identity. Previous methods either required reference utterances during inference, did not preserve speaker identity well, or used one-to-one systems that could only be trained for each non-native accent. This paper presents a promising AC model that can convert many accents into native to overcome these issues. Our approach utilizes discrete units, derived from clustering self-supervised representations of native speech, as an intermediary target for accent conversion. Leveraging multi-speaker text-to-speech synthesis, it transforms these discrete representations back into native speech while retaining the speaker identity. Additionally, we develop an efficient data augmentation method to train the system without demanding a lot of non-native resources. Our system is proved to improve non-native speaker fluency, sound like a native accent, and preserve original speaker identity well.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。