arXiv:2409.11107eess.AScs.SD2024-09中稿 · the Asilomar 2023 …被引 3

用零样本语音合成扩充低资源口音数据,提升自动语音识别性能

Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora

  • 用零样本文本转语音生成口音语音,补充真实数据不足
  • 口音语音识别错误率降低5%,最高达14%改善
  • 适合低资源口音数据的ASR系统优化,无需大量标注

近年来,自动语音识别(ASR)模型在清晰、低噪和混响环境下的识别性能显著提升。然而,这些系统均依赖特定声学环境下数百小时的标注数据。当此类数据缺失时,系统性能会大幅下降,尤其在特定声学环境或说话人群体未充分代表时。本文研究了口音语音对现成ASR系统的影响,并提出一种基于零样本文本到语音(TTS)的语料增强策略。实验表明,该方法可将口音语音上的识别错误率降低最多5%(即相对减少5%),且仅使用少量真实数据与合成数据混合训练,其性能优于纯真实口音数据训练模型,最高实现14%的词错误率(WERR)降低。

原文摘要 · Abstract (English)

In recent years, automatic speech recognition (ASR) models greatly improved transcription performance both in clean, low noise, acoustic conditions and in reverberant environments. However, all these systems rely on the availability of hundreds of hours of labelled training data in specific acoustic conditions. When such a training dataset is not available, the performance of the system is heavily impacted. For example, this happens when a specific acoustic environment or a particular population of speakers is under-represented in the training dataset. Specifically, in this paper we investigate the effect of accented speech data on an off-the-shelf ASR system. Furthermore, we suggest a strategy based on zero-shot text-to-speech to augment the accented speech corpora. We show that this augmentation method is able to mitigate the loss in performance of the ASR system on accented data up to 5% word error rate reduction (WERR). In conclusion, we demonstrate that by incorporating a modest fraction of real with synthetically generated data, the ASR system exhibits superior performance compared to a model trained exclusively on authentic accented speech with up to 14% WERR.

语音识别口音处理数据增强零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。