通过内部声学模型提升语音识别效率,实现42%-75%加速。
Boosting Hybrid Autoregressive Transducer-based ASR with Internal Acoustic Model Training and Dual Blank Thresholding
- 引入共享网络的内部声学模型,与主模型联合训练。
- 结合双空白阈值策略,解码速度提升42%-75%且性能稳定。
- 适合追求高效实时语音识别的工程应用者。
混合自回归转换器(HAT)是一种将空白与非空白后验分布分开建模的神经转换器变体。本文提出一种新型内部声学模型(IAM)训练策略,以增强基于HAT的语音识别性能。IAM由编码器和联合网络组成,与HAT完全共享并联合训练。该联合训练不仅提升了HAT的训练效率,还促使IAM与HAT同步输出空白符号,从而跳过更耗时的非空白计算,实现更高效的空白阈值控制,加快解码速度。实验表明,相比原始HAT,加入IAM的模型在相对错误率上显著降低。此外,我们引入双空白阈值机制,融合HAT与IAM的空白阈值,并设计兼容解码算法。该方法在不造成明显性能损失的前提下,实现42%-75%的解码速度提升。
原文摘要 · Abstract (English)
A hybrid autoregressive transducer (HAT) is a variant of neural transducer that models blank and non-blank posterior distributions separately. In this paper, we propose a novel internal acoustic model (IAM) training strategy to enhance HAT-based speech recognition. IAM consists of encoder and joint networks, which are fully shared and jointly trained with HAT. This joint training not only enhances the HAT training efficiency but also encourages IAM and HAT to emit blanks synchronously which skips the more expensive non-blank computation, resulting in more effective blank thresholding for faster decoding. Experiments demonstrate that the relative error reductions of the HAT with IAM compared to the vanilla HAT are statistically significant. Moreover, we introduce dual blank thresholding, which combines both HAT- and IAM-blank thresholding and a compatible decoding algorithm. This results in a 42-75% decoding speed-up with no major performance degradation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。