arXiv:2506.23367cs.SDcs.CL2025-06中稿 · ISCA Speech Synthe…

用元音时长差异提升英语二语发音清晰度,让听者更易懂且感觉更尊重。

You Sound a Little Tense: L2 Tailored Clear TTS Using Durational Vowel Properties

  • 利用美式英语紧元音与松元音的时长差异设计清晰模式
  • 法语母语者在清晰模式下错误率降低至少9.15%
  • 适合关注二语语音可懂性与用户体验的研究者

我们提出了首个面向第二语言(L2)使用者的文本到语音(TTS)系统。通过利用美式英语中紧元音(较长)与松元音(较短)之间的时长差异,为Matcha-TTS设计了“清晰模式”。感知实验表明,以法语为母语、英语为第二语言的听者在使用该清晰模式时,转录错误率至少降低了9.15%,且认为该模式更具鼓励性与尊重感,优于整体减速的语音。令人惊讶的是,听者并未察觉这些效果:尽管清晰模式显著降低错误率,他们仍认为整体减速最易懂,说明实际可懂性与感知可懂性并不相关。此外,我们发现Whisper-ASR并未使用与二语学习者相同的线索来区分难辨元音,因此不足以评估此类人群的语音可懂性。

原文摘要 · Abstract (English)

We present the first text-to-speech (TTS) system tailored to second language (L2) speakers. We use duration differences between American English tense (longer) and lax (shorter) vowels to create a "clarity mode" for Matcha-TTS. Our perception studies showed that French-L1, English-L2 listeners had fewer (at least 9.15%) transcription errors when using our clarity mode, and found it more encouraging and respectful than overall slowed down speech. Remarkably, listeners were not aware of these effects: despite the decreased word error rate in clarity mode, listeners still believed that slowing all target words was the most intelligible, suggesting that actual intelligibility does not correlate with perceived intelligibility. Additionally, we found that Whisper-ASR did not use the same cues as L2 speakers to differentiate difficult vowels and is not sufficient to assess the intelligibility of TTS systems for these individuals.

语音合成二语语音清晰度优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。