用声学后验图实现芬兰语发音编辑,无需文本对齐。
Pronunciation Editing for Finnish Speech using Phonetic Posteriorgrams
- 基于扩散模型的音素后验图转语音,支持单音素编辑。
- 在约60小时数据上实现自然度与发音修正效果双提升。
- 适合低资源语言语音合成与语言学习场景研究者。
合成第二语言(L2)语音对语言学习体验和反馈具有潜在重要价值。然而,由于缺乏L2语音合成数据集,低资源语言的L2语音合成面临挑战。本文提出一种实用方法,通过编辑母语语音来近似L2语音,并介绍PPG2Speech——一种基于扩散模型的多说话人音素后验图到语音(PPG2Speech)模型,可实现无文本对齐的单音素编辑。该模型以Matcha-TTS的流匹配解码器为骨干,将音素后验图(PPGs)转换为梅尔频谱图,条件依赖于外部说话人嵌入和音高。通过引入无分类器引导(CFG)和摆动采样(Sway Sampling),强化了原有解码器能力。我们还提出了一个新的任务特定客观评估指标——音素对齐一致性(PAC),用于衡量编辑后的PPGs与合成语音提取的PPGs之间的一致性。在芬兰语(一种低资源、近乎音素化语言)上进行了验证,使用约60小时数据。通过客观与主观评估,对比了本方法在自然度、说话人相似性及编辑有效性方面与基于TTS的编辑方法的表现。代码已公开于https://github.com/aalto-speech/PPG2Speech。
原文摘要 · Abstract (English)
Synthesizing second-language (L2) speech is potentially highly valued for L2 language learning experience and feedback. However, due to the lack of L2 speech synthesis datasets, it is difficult to synthesize L2 speech for low-resourced languages. In this paper, we provide a practical solution for editing native speech to approximate L2 speech and present PPG2Speech, a diffusion-based multispeaker Phonetic-Posteriorgrams-to-Speech model that is capable of editing a single phoneme without text alignment. We use Matcha-TTS's flow-matching decoder as the backbone, transforming Phonetic Posteriorgrams (PPGs) to mel-spectrograms conditioned on external speaker embeddings and pitch. PPG2Speech strengthens the Matcha-TTS's flow-matching decoder with Classifier-free Guidance (CFG) and Sway Sampling. We also propose a new task-specific objective evaluation metric, the Phonetic Aligned Consistency (PAC), between the edited PPGs and the PPGs extracted from the synthetic speech for editing effects. We validate the effectiveness of our method on Finnish, a low-resourced, nearly phonetic language, using approximately 60 hours of data. We conduct objective and subjective evaluations of our approach to compare its naturalness, speaker similarity, and editing effectiveness with TTS-based editing. Our source code is published at https://github.com/aalto-speech/PPG2Speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。