标注语调分段能提升口语合成自然度,人工标注更接近真实语调。
The Impact of Prosodic Segmentation on Speech Synthesis of Spontaneous Speech
- 对比人工与自动语调分段对语音合成的影响
- 人工标注使合成语音更自然,识别率略有提升
- 适合研究口语合成与语调建模的学者使用
自发性口语在语音合成中面临诸多挑战,如对话中的换句、停顿和不流畅现象。尽管当前语音合成系统在生成自然、可懂语音方面取得显著进展,主要依赖隐式建模语调特征(如音高、强度、时长)的架构,但显式标注语调分段的数据集构建及其对自发性口语合成的影响仍鲜有研究。本文评估了在巴西葡萄牙语中,人工与自动语调分段标注对非自回归模型FastSpeech 2合成质量的影响。实验结果表明,使用语调分段训练可略微提升语音的可懂性和声学自然度。虽然自动标注产生的分段更规整,但人工标注引入更大变异,有助于实现更自然的语调。对中性陈述句的分析显示,两种训练方式均重现了预期的重音模式,但语调模型更贴近自然的前重音轮廓。为支持可复现性与未来研究,所有数据集、源代码及训练模型均已公开,采用CC BY-NC-ND 4.0许可。
原文摘要 · Abstract (English)
Spontaneous speech presents several challenges for speech synthesis, particularly in capturing the natural flow of conversation, including turn-taking, pauses, and disfluencies. Although speech synthesis systems have made significant progress in generating natural and intelligible speech, primarily through architectures that implicitly model prosodic features such as pitch, intensity, and duration, the construction of datasets with explicit prosodic segmentation and their impact on spontaneous speech synthesis remains largely unexplored. This paper evaluates the effects of manual and automatic prosodic segmentation annotations in Brazilian Portuguese on the quality of speech synthesized by a non-autoregressive model, FastSpeech 2. Experimental results show that training with prosodic segmentation produced slightly more intelligible and acoustically natural speech. While automatic segmentation tends to create more regular segments, manual prosodic segmentation introduces greater variability, which contributes to more natural prosody. Analysis of neutral declarative utterances showed that both training approaches reproduced the expected nuclear accent pattern, but the prosodic model aligned more closely with natural pre-nuclear contours. To support reproducibility and future research, all datasets, source codes, and trained models are publicly available under the CC BY-NC-ND 4.0 license.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。