用语音转换技术实现隆巴德语调风格迁移,提升语音可懂度。
Voice Conversion for Lombard Speaking Style with Implicit and Explicit Acoustic Feature Conditioning
- 通过隐式声学特征条件控制,实现风格保留的语音转换。
- 隐式方法在可懂度上达到与显式方法相当的效果。
- 适合语音增强、助听及嘈杂环境下的语音合成研究者。
在隆巴德语调下,文本转语音(TTS)系统能显著提升语音可懂度,对听力障碍者和嘈杂环境具有重要意义。然而,训练此类模型需大量数据,而隆巴德效应因说话人差异、噪声变化及录制疲劳难以有效采集。语音转换(VC)被证明是一种有效数据增强手段,可在缺乏目标说话人特定语调录音时训练TTS系统。本文聚焦隆巴德语调风格迁移,目标是在变换说话人身份的同时,保留定义该语调的声学特征。我们对比了隐式与显式声学特征条件下的语音转换模型,发现所提出的隐式条件策略在可懂度上与显式条件模型相当,同时保持了更好的说话人相似性。
原文摘要 · Abstract (English)
Text-to-Speech (TTS) systems in Lombard speaking style can improve the overall intelligibility of speech, useful for hearing loss and noisy conditions. However, training those models requires a large amount of data and the Lombard effect is challenging to record due to speaker and noise variability and tiring recording conditions. Voice conversion (VC) has been shown to be a useful augmentation technique to train TTS systems in the absence of recorded data from the target speaker in the target speaking style. In this paper, we are concerned with Lombard speaking style transfer. Our goal is to convert speaker identity while preserving the acoustic attributes that define the Lombard speaking style. We compare voice conversion models with implicit and explicit acoustic feature conditioning. We observe that our proposed implicit conditioning strategy achieves an intelligibility gain comparable to the model conditioned on explicit acoustic features, while also preserving speaker similarity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。