让语音合成具备像人一样在嘈杂环境变大声变清晰的能力。
Synthesizing the Lombard Effect: Multi-Level Control of Speech Clarity and Vocal Effort in TTS

- 基于流匹配的语音合成模型,可连续控制发声力度与发音清晰度。
- 在噪声环境中合成语音的可懂度提升,接近真人清晰说话效果。
- 支持词级强调,适合需要突出重点语句的应用场景。
人类在嘈杂环境或面对听力障碍者时会不自觉地提高音量并增强发音清晰度,这种现象称为隆巴德效应(Lombard effect)。为在语音合成系统中模拟该行为,本文提出一种基于流匹配的文本转语音(TTS)模型,利用发声力度与发音清晰度伪标签进行训练。该模型实现了发声力度与发音清晰度的连续且解耦控制,同时支持词级语义强调,以增强特定语段的可理解性。实验表明,该控制机制显著提升了与清晰度相关的声学特征。此外,语音-噪声环境下测试显示,该模型成功模拟了人类清晰语音在噪声中的可懂度增益。
原文摘要 · Abstract (English)
Humans tend to speak louder and clearer in challenging environments, such as noisy conditions or when addressing hearingimpaired listeners, which is called Lombard effect. To simulate this behavior in speech synthesis systems, we introduce a flow-matching based text-to-speech (TTS) model trained with vocal effort and articulation pseudo-labels. The proposed model achieves continuous and disentangled control of vocal effort and articulation, while also enabling word-level emphasis for clarifying specific segments of an utterance. Experimental results show that these control mechanisms effectively improve clarityrelated acoustic features. Furthermore, speech-in-noise experiments demonstrate that our model successfully simulates the intelligibility gains of human clear speech in noisy conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。