arXiv:2409.12745cs.SDcs.AI2024-09

用ASR过滤提升语音合成数据质量,改善语音命令识别效果

Enhancing Synthetic Training Data for Speech Commands: From ASR-Based Filtering to Domain Adaptation in SSL Latent Space

  • 用ASR检测筛选合成语音,去除低质量样本
  • 经过滤后模型准确率显著提升,达94.2%
  • 合成与真实语音在WavLM特征空间仍可区分,适合语音增强研究者

合成语音作为数据增强手段在自动语音识别和语音分类任务中日益流行。尽管具备语音克隆能力的新型文本转语音系统可基于短音频片段生成更多声音,但这些系统常产生幻觉,生成不良数据,对下游任务造成负面影响。本文针对语音命令分类任务,在Google Speech Commands数据集上开展零样本学习实验。结果表明,简单的ASR基过滤方法能显著提升生成数据质量,进而提升模型性能。此外,即使生成语音质量良好,其在自监督模型WavLM特征空间中仍与真实语音明显可区分,该问题通过CycleGAN进行域适应进一步缓解。

原文摘要 · Abstract (English)

The use of synthetic speech as data augmentation is gaining increasing popularity in fields such as automatic speech recognition and speech classification tasks. Despite novel text-to-speech systems with voice cloning capabilities, that allow the usage of a larger amount of voices based on short audio segments, it is known that these systems tend to hallucinate and oftentimes produce bad data that will most likely have a negative impact on the downstream task. In the present work, we conduct a set of experiments around zero-shot learning with synthetic speech data for the specific task of speech commands classification. Our results on the Google Speech Commands dataset show that a simple ASR-based filtering method can have a big impact in the quality of the generated data, translating to a better performance. Furthermore, despite the good quality of the generated speech data, we also show that synthetic and real speech can still be easily distinguishable when using self-supervised (WavLM) features, an aspect further explored with a CycleGAN to bridge the gap between the two types of speech material.

语音合成数据增强自监督学习语音分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。