无需训练数据,即可生成任意人不同强度的嘈杂环境说话声。
Lombard Speech Synthesis for Any Voice with Controllable Style Embeddings
- 用主成分分析找到影响发声风格的关键特征向量
- 通过调整特征向量实现对语调的精细控制
- 适合需要自定义说话风格的语音合成应用
Lombard效应在嘈杂环境或面对听力障碍者时对自然交流至关重要。我们提出一种可控文本到语音(TTS)系统,可在不使用显式Lombard数据训练的情况下,为任意说话人合成Lombard语音。该方法利用大规模韵律多样数据集学习风格嵌入,并通过主成分分析(PCA)分析其与Lombard属性的相关性。通过移动相关主成分,调整风格嵌入并融入TTS模型,实现对目标Lombard水平的语音生成。评估表明,该方法保持了语音自然度和说话人身份,提升了噪声下的可懂度,并实现了对韵律的细粒度控制,为任意说话人提供了鲁棒的可控Lombard TTS解决方案。
原文摘要 · Abstract (English)
The Lombard effect plays a key role in natural communication, particularly in noisy environments or when addressing hearing-impaired listeners. We present a controllable text-to-speech (TTS) system capable of synthesizing Lombard speech for any speaker without requiring explicit Lombard data during training. Our approach leverages style embeddings learned from a large, prosodically diverse dataset and analyzes their correlation with Lombard attributes using principal component analysis (PCA). By shifting the relevant PCA components, we manipulate the style embeddings and incorporate them into our TTS model to generate speech at desired Lombard levels. Evaluations demonstrate that our method preserves naturalness and speaker identity, enhances intelligibility under noise, and provides fine-grained control over prosody, offering a robust solution for controllable Lombard TTS for any speaker.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。