用多种音频表示训练模型,再融合提升语音识别效果
Late fusion ensembles for speech recognition on diverse input audio representations
- 不同音频表示输入的E-Branchformer模型做后期融合
- 在4个数据集上提升1%~14%,超越当前最佳模型
- 即使有语言模型也有效,适合追求高精度的ASR研究者
本文探索了语音音频的多种表示形式及其对晚期融合集成的E-Branchformer模型在自动语音识别(ASR)任务中的影响。尽管集成方法通常能提升系统性能,但本研究关注的是中型和大型E-Branchformer这类复杂先进模型,在基于不同输入音频表示训练时,其集成表现如何。实验在四个广泛使用的基准数据集(Librispeech、Aishell、Gigaspeech、TEDLIUMv2)上进行,结果表明,与采用相似技术训练的现有最优模型相比,仍可实现1%至14%的性能提升。值得注意的是,即便使用语言模型,这种集成仍能带来改进,尽管差距正在缩小。
原文摘要 · Abstract (English)
We explore diverse representations of speech audio, and their effect on a performance of late fusion ensemble of E-Branchformer models, applied to Automatic Speech Recognition (ASR) task. Although it is generally known that ensemble methods often improve the performance of the system even for speech recognition, it is very interesting to explore how ensembles of complex state-of-the-art models, such as medium-sized and large E-Branchformers, cope in this setting when their base models are trained on diverse representations of the input speech audio. The results are evaluated on four widely-used benchmark datasets: \textit{Librispeech, Aishell, Gigaspeech}, \textit{TEDLIUMv2} and show that improvements of $1\% - 14\%$ can still be achieved over the state-of-the-art models trained using comparable techniques on these datasets. A noteworthy observation is that such ensemble offers improvements even with the use of language models, although the gap is closing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。