仅用母语数据,通过离散令牌重构语音,实现更真实的外语口音模拟。
Prosodically Enhanced Foreign Accent Simulation by Discrete Token-based Resynthesis Only with Native Speech Corpora
- 基于自监督学习的离散令牌,仅用母语数据重构语音
- 成功复现第二语言者特有的时长口音特征
- 适合语音识别鲁棒性训练与跨语言语音研究
近期提出一种仅使用母语语音数据,通过自监督学习(SSL)模型获得的离散令牌来合成外语口音语音的方法。鉴于带口音语音数据稀缺,该方法极大降低了外语口音模拟的难度。利用合成的口音语音作为人类听觉材料或自动语音识别(ASR)的训练数据,可提升系统对外国口音的鲁棒性。然而,此前方法存在致命缺陷:无法再现与时长相关的口音。这类口音常见于母语为音节时序或音拍时序的语言者(如日语)说重音时序语言(如英语)时。本文在原有方法中引入时长调节机制,使口音模拟更加准确。实验表明,所提方法能有效复现真实第二语言语音中的时长口音特征。
原文摘要 · Abstract (English)
Recently, a method for synthesizing foreign-accented speech only with native speech data using discrete tokens obtained from self-supervised learning (SSL) models was proposed. Considering limited availability of accented speech data, this method is expected to make it much easier to simulate foreign accents. By using the synthesized accented speech as listening materials for humans or training data for automatic speech recognition (ASR), both of them will acquire higher robustness against foreign accents. However, the previous method has a fatal flaw that it cannot reproduce duration-related accents. Durational accents are commonly seen when L2 speakers, whose native language has syllable-timed or mora-timed rhythm, speak stress-timed languages, such as English. In this paper, we integrate duration modification to the previous method to simulate foreign accents more accurately. Experiments show that the proposed method successfully replicates durational accents seen in real L2 speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。