arXiv:2409.07151eess.AScs.AI2024-09被引 2

用零样本语音合成生成学习者专属标准发音,提升二语发音评估效果

Zero-Shot Text-to-Speech as Golden Speech Generator: A Systematic Framework and its Applicability in Automatic Pronunciation Assessment

  • 基于零样本语音合成生成个性化标准发音作为评估基准
  • 在L2-ARCTIC和Speechocean762数据集上显著提升评估性能
  • 为自动发音训练系统提供新范式,适合语言教育技术研究者

第二语言学习者可通过模仿符合自身语音特征的标准发音来改善发音。本研究探索假设:利用零样本文本转语音(ZS-TTS)技术生成的学习者专属标准发音,可作为衡量二语发音水平的有效指标。研究贡献有两点:一是设计并开发了评估合成模型生成标准发音能力的系统性框架;二是深入探究了使用标准发音进行自动发音评估(APA)的有效性。在L2-ARCTIC和Speechocean762基准数据集上的全面实验表明,所提方法在多种评估指标上均显著优于现有方法。据我们所知,这是首次系统探讨标准发音在ZS-TTS与APA中作用的研究,为计算机辅助发音训练(CAPT)提供了有前景的新范式。

原文摘要 · Abstract (English)

Second language (L2) learners can improve their pronunciation by imitating golden speech, especially when the speech that aligns with their respective speech characteristics. This study explores the hypothesis that learner-specific golden speech generated with zero-shot text-to-speech (ZS-TTS) techniques can be harnessed as an effective metric for measuring the pronunciation proficiency of L2 learners. Building on this exploration, the contributions of this study are at least two-fold: 1) design and development of a systematic framework for assessing the ability of a synthesis model to generate golden speech, and 2) in-depth investigations of the effectiveness of using golden speech in automatic pronunciation assessment (APA). Comprehensive experiments conducted on the L2-ARCTIC and Speechocean762 benchmark datasets suggest that our proposed modeling can yield significant performance improvements with respect to various assessment metrics in relation to some prior arts. To our knowledge, this study is the first to explore the role of golden speech in both ZS-TTS and APA, offering a promising regime for computer-assisted pronunciation training (CAPT).

语音合成发音评估二语学习零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。