用生成模型模拟不同录音环境,提升语音识别在陌生场景下的鲁棒性。
Channel-Aware Domain-Adaptive Generative Adversarial Network for Robust Speech Recognition
- 通过提取音频通道特征,引导GAN合成匹配目标环境的语音。
- 在客家话和台湾话数据集上,字符错误率降低20.02%和9.64%。
- 适合需要跨环境语音识别的工业应用,如智能客服、语音助手。
预训练语音识别系统在匹配语境下表现优异,但在未见录音环境导致的通道不匹配时性能显著下降。为此,我们提出一种新型通道感知的数据仿真方法,用于增强鲁棒性语音识别训练。该方法结合通道提取技术和生成对抗网络(GAN),首先训练一个可从任意音频中提取通道嵌入的编码器;随后利用少量目标域数据提取通道嵌入,并以此指导基于GAN的语音合成器。该合成器在保持输入语音音素内容不变的同时,模拟目标域的通道特性。我们在具有挑战性的客家话跨台湾(HAT)和台湾话跨台湾(TAT)语料库上评估该方法,相对基线分别实现20.02%和9.64%的字符错误率(CER)降低。结果表明,该通道感知数据仿真方法能有效弥合源域与目标域声学特征的差距。
原文摘要 · Abstract (English)
While pre-trained automatic speech recognition (ASR) systems demonstrate impressive performance on matched domains, their performance often degrades when confronted with channel mismatch stemming from unseen recording environments and conditions. To mitigate this issue, we propose a novel channel-aware data simulation method for robust ASR training. Our method harnesses the synergistic power of channel-extractive techniques and generative adversarial networks (GANs). We first train a channel encoder capable of extracting embeddings from arbitrary audio. On top of this, channel embeddings are extracted using a minimal amount of target-domain data and used to guide a GAN-based speech synthesizer. This synthesizer generates speech that faithfully preserves the phonetic content of the input while mimicking the channel characteristics of the target domain. We evaluate our method on the challenging Hakka Across Taiwan (HAT) and Taiwanese Across Taiwan (TAT) corpora, achieving relative character error rate (CER) reductions of 20.02% and 9.64%, respectively, compared to the baselines. These results highlight the efficacy of our channel-aware data simulation method for bridging the gap between source- and target-domain acoustics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。