AMNet提升普通话语音合成质量,通过分句解析与局部卷积增强语调表现。
AMNet: An Acoustic Model Network for Enhanced Mandarin Speech Synthesis
- 引入分句解析与局部卷积模块,增强对停顿、重音等局部语音特征的建模能力。
- 在主观评测中获更高平均意见分(MOS),MCD降低,基频拟合相关系数达0.92。
- 特别适合需要高保真语调表达的中文语音合成应用,如有声书、虚拟助手。
本文提出AMNet,一种用于提升普通话语音合成性能的声学模型网络,通过融合分句结构标注和局部卷积模块实现。AMNet基于FastSpeech 2架构,针对局部上下文建模难题进行优化,该问题对捕捉停顿、重音和语调等复杂语音特征至关重要。模型嵌入分句解析器并引入局部卷积模块,增强对局部信息的敏感度。同时,将声调特征与音素解耦,提供显式声调建模指导,提升声调准确性和发音质量。实验表明,相比基线模型,AMNet在主观与客观评估中均表现更优:取得更高的平均意见分(MOS),更低的梅尔倒谱失真(MCD),以及更好的基频拟合效果(F0 R² = 0.92),证实其能生成高质量、自然且富有表现力的普通话语音。
原文摘要 · Abstract (English)
This paper presents AMNet, an Acoustic Model Network designed to improve the performance of Mandarin speech synthesis by incorporating phrase structure annotation and local convolution modules. AMNet builds upon the FastSpeech 2 architecture while addressing the challenge of local context modeling, which is crucial for capturing intricate speech features such as pauses, stress, and intonation. By embedding a phrase structure parser into the model and introducing a local convolution module, AMNet enhances the model's sensitivity to local information. Additionally, AMNet decouples tonal characteristics from phonemes, providing explicit guidance for tone modeling, which improves tone accuracy and pronunciation. Experimental results demonstrate that AMNet outperforms baseline models in subjective and objective evaluations. The proposed model achieves superior Mean Opinion Scores (MOS), lower Mel Cepstral Distortion (MCD), and improved fundamental frequency fitting $F0 (R^2)$, confirming its ability to generate high-quality, natural, and expressive Mandarin speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。