arXiv:2411.17537eess.AScs.LG2024-11被引 1

提出新方法提升流式语音识别模型训练精度

Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition

  • 引入前向变量因果补偿机制,量化训练与推理的差异
  • 在LibriSpeech上实验显示,新方法显著提升流式识别准确率
  • 适合关注实时语音识别模型优化的研究者

基于转换器的神经网络已成为流式自动语音识别(ASR)的主流方法,在准确率与延迟之间取得了最优平衡。传统框架下,流式转换器模型采用非流式递归规则最大化似然函数进行训练,导致训练与推理不匹配,引发似然函数畸变,进而影响识别准确率。本文提出了前向变量因果补偿(FoCC),对实际似然与畸变似然之间的差距进行数学量化,并设计了其估计器FoCCE,以实现精确似然估计。在LibriSpeech数据集上的实验表明,采用FoCCE训练可有效提升流式转换器的识别性能。

原文摘要 · Abstract (English)

Transducer neural networks have emerged as the mainstream approach for streaming automatic speech recognition (ASR), offering state-of-the-art performance in balancing accuracy and latency. In the conventional framework, streaming transducer models are trained to maximize the likelihood function based on non-streaming recursion rules. However, this approach leads to a mismatch between training and inference, resulting in the issue of deformed likelihood and consequently suboptimal ASR accuracy. We introduce a mathematical quantification of the gap between the actual likelihood and the deformed likelihood, namely forward variable causal compensation (FoCC). We also present its estimator, FoCCE, as a solution to estimate the exact likelihood. Through experiments on the LibriSpeech dataset, we show that FoCCE training improves the accuracy of the streaming transducers.

语音识别流式处理似然优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。