研究说话人嵌入与发音规则的交互,提升语音合成的方言控制力。
Quantifying Speaker Embedding Phonological Rule Interactions in Accented Speech Synthesis
- 用发音规则建模英美口音差异,结合说话人嵌入生成真实口音。
- 提出新指标PSR,量化嵌入对规则的抑制或覆盖程度。
- 揭示口音与说话人身份的耦合关系,适合语音合成研究者参考。
许多语言如英语存在广泛方言和口音差异,使口音控制成为灵活文本转语音(TTS)模型的关键能力。当前TTS系统通常通过关联特定口音的说话人嵌入来生成带口音语音,虽有效但解释性与可控性有限,因嵌入同时编码音色、情感等属性。本文以美式与英式英语为例,实现闪音、卷舌化及元音对应等发音规则。提出音素偏移率(PSR)作为新指标,量化嵌入对规则转换的保留或覆盖强度。实验表明,规则与嵌入结合可生成更自然的口音,但嵌入会削弱或覆盖规则,揭示口音与说话人身份之间的纠缠。研究强调规则是口音控制的有效杠杆,并提供评估语音生成解耦性的框架。
原文摘要 · Abstract (English)
Many spoken languages, including English, exhibit wide variation in dialects and accents, making accent control an important capability for flexible text-to-speech (TTS) models. Current TTS systems typically generate accented speech by conditioning on speaker embeddings associated with specific accents. While effective, this approach offers limited interpretability and controllability, as embeddings also encode traits such as timbre and emotion. In this study, we analyze the interaction between speaker embeddings and linguistically motivated phonological rules in accented speech synthesis. Using American and British English as a case study, we implement rules for flapping, rhoticity, and vowel correspondences. We propose the phoneme shift rate (PSR), a novel metric quantifying how strongly embeddings preserve or override rule-based transformations. Experiments show that combining rules with embeddings yields more authentic accents, while embeddings can attenuate or overwrite rules, revealing entanglement between accent and speaker identity. Our findings highlight rules as a lever for accent control and a framework for evaluating disentanglement in speech generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。