arXiv:2409.11742cs.SDeess.AS2024-09

用语音转换模拟母语者跟读,评估非母语发音清晰度。

Simulating Native Speaker Shadowing for Nonnative Speech Assessment with Latent Speech Representations

  • 通过隐式语音表征与语音转换技术模拟母语者跟读过程。
  • 自监督语音表征使生成结果在语言准确性和自然度上更接近真实跟读。
  • 适合语言学习系统评估发音质量,无需真人参与。

评估语音可理解性是计算机辅助语言学习系统中的关键任务。传统方法通常依赖自动语音识别(ASR)提供的词错误率(WER)作为可理解性评分,但该方法因人类语音识别(HSR)与ASR之间存在显著差异而受限。一个有前景的替代方案是让母语者(L1)对非母语者(L2)的发言进行跟读。若母语者在跟读中出现中断或误读,则可作为评估L2语音可理解性的指标。本研究提出一种语音生成系统,利用语音转换(VC)技术与隐式语音表征模拟母语者跟读过程。实验结果表明,该方法能有效复现母语者跟读行为,为评估L2语音可理解性提供新工具。值得注意的是,采用自监督语音表征(S3R)的系统在语言准确性和自然度方面与真实母语者跟读更为相似。

原文摘要 · Abstract (English)

Evaluating speech intelligibility is a critical task in computer-aided language learning systems. Traditional methods often rely on word error rates (WER) provided by automatic speech recognition (ASR) as intelligibility scores. However, this approach has significant limitations due to notable differences between human speech recognition (HSR) and ASR. A promising alternative is to involve a native (L1) speaker in shadowing what nonnative (L2) speakers say. Breakdowns or mispronunciations in the L1 speaker's shadowing utterance can serve as indicators for assessing L2 speech intelligibility. In this study, we propose a speech generation system that simulates the L1 shadowing process using voice conversion (VC) techniques and latent speech representations. Our experimental results demonstrate that this method effectively replicates the L1 shadowing process, offering an innovative tool to evaluate L2 speech intelligibility. Notably, systems that utilize self-supervised speech representations (S3R) show a higher degree of similarity to real L1 shadowing utterances in both linguistic accuracy and naturalness.

语音评估语音转换自监督学习语言学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。