arXiv:2411.13159cs.CLcs.SD2024-11被引 2

用大模型生成难识别语音,提升语音识别准确率

Hard-Synth: Synthesizing Diverse Hard Samples for ASR using Zero-Shot TTS and LLM

  • 用大模型重写文本并零样本克隆难识别语音风格
  • 在LibriSpeech上使词错误率降低6.5%和4.4%
  • 适合需要数据高效和降低偏见的语音识别场景

文本到语音(TTS)模型被广泛用于仅基于文本语料增强自动语音识别(ASR)系统,从而降低真实语音标注数据的成本。现有研究主要依赖额外文本数据和预定义语音风格。本文提出Hard-Synth,一种新型ASR数据增强方法,结合大语言模型(LLM)与先进零样本TTS。该方法利用LLM生成多样化领域内文本,无需额外文本数据;不采用预设语音风格,而是通过硬提示选择法,使用零样本TTS克隆ASR模型难以识别的语音风格。实验表明,Hard-Synth显著提升了Conformer模型性能,在LibriSpeech dev/test-other子集上分别实现6.5%和4.4%的相对词错误率(WER)下降。此外,Hard-Synth具有数据高效性,可有效减少ASR中的偏差。

原文摘要 · Abstract (English)

Text-to-speech (TTS) models have been widely adopted to enhance automatic speech recognition (ASR) systems using text-only corpora, thereby reducing the cost of labeling real speech data. Existing research primarily utilizes additional text data and predefined speech styles supported by TTS models. In this paper, we propose Hard-Synth, a novel ASR data augmentation method that leverages large language models (LLMs) and advanced zero-shot TTS. Our approach employs LLMs to generate diverse in-domain text through rewriting, without relying on additional text data. Rather than using predefined speech styles, we introduce a hard prompt selection method with zero-shot TTS to clone speech styles that the ASR model finds challenging to recognize. Experiments demonstrate that Hard-Synth significantly enhances the Conformer model, achieving relative word error rate (WER) reductions of 6.5\%/4.4\% on LibriSpeech dev/test-other subsets. Additionally, we show that Hard-Synth is data-efficient and capable of reducing bias in ASR.

语音识别数据增强零样本TTS大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。