让CTC模型显式学习上下文依赖的内部语言模型,提升跨领域语音识别效果。
Label-Context-Dependent Internal Language Model Estimation for CTC
- 通过知识蒸馏显式建模CTC的上下文依赖语言模型
- 跨域评估中词错误率降低超13%(相对)
- 适合提升语音识别在陌生语境下的表现
尽管连接时序分类(CTC)假设标签上下文独立,但现代强大编码器仍使其隐式学习到上下文依赖的内部语言模型(ILM)。本文研究了CTC中隐含的上下文依赖性,提出基于知识蒸馏(KD)的新型上下文依赖ILM估计方法,并给出理论依据。同时引入两种KD正则化方法。在Librispeech和TED-LIUM Release 2数据集上进行域内与跨域评估。实验表明,上下文依赖的ILM在跨域评估中优于上下文独立先验,证明了CTC确实学习了上下文依赖的ILM。所提标签级知识蒸馏结合平滑方法,在词错误率上相比浅融合提升超过13%的相对性能。
原文摘要 · Abstract (English)
Although connectionist temporal classification (CTC) has the label context independence assumption, it can still implicitly learn a context-dependent internal language model (ILM) due to modern powerful encoders. In this work, we investigate the implicit context dependency modeled in the ILM of CTC. To this end, we propose novel context-dependent ILM estimation methods for CTC based on knowledge distillation (KD) with theoretical justifications. Furthermore, we introduce two regularization methods for KD. We conduct experiments on Librispeech and TED-LIUM Release 2 datasets for in-domain and cross-domain evaluation, respectively. Experimental results show that context-dependent ILMs outperform the context-independent priors in cross-domain evaluation, indicating that CTC learns a context-dependent ILM. The proposed label-level KD with smoothing method surpasses other ILM estimation approaches, with more than 13% relative improvement in word error rate compared to shallow fusion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。