提出联合无监督与有监督训练框架,提升语音识别模型泛化与任务适配性。
Bilevel Joint Unsupervised and Supervised Training for Automatic Speech Recognition
- 双层优化同时最小化无监督与有监督损失,实现端到端联合训练
- 在多个数据集上优于传统预训练+微调及主流半监督方法
- 适合资源有限场景下提升语音识别模型性能,尤其适用于标注数据少的场景
本文提出一种双层联合无监督与有监督训练(BL-JUST)框架用于自动语音识别。相较于传统的预训练-微调两阶段分离流程,BL-JUST通过优化声学模型,使其同时最小化无监督和有监督损失函数。由于该方法寻求两种损失函数的匹配局部最优解,所学习到的声学表示在通用性和任务特定性之间取得良好平衡。我们采用基于惩罚项的双层梯度下降法求解该问题,并在多种数据集、网络架构和损失函数配置下评估了训练得到的深度神经网络声学模型。实验表明,BL-JUST可超越广泛使用的预训练-微调策略及其他主流半监督技术。
原文摘要 · Abstract (English)
In this paper, we propose a bilevel joint unsupervised and supervised training (BL-JUST) framework for automatic speech recognition. Compared to the conventional pre-training and fine-tuning strategy which is a disconnected two-stage process, BL-JUST tries to optimize an acoustic model such that it simultaneously minimizes both the unsupervised and supervised loss functions. Because BL-JUST seeks matched local optima of both loss functions, acoustic representations learned by the acoustic model strike a good balance between being generic and task-specific. We solve the BL-JUST problem using penalty-based bilevel gradient descent and evaluate the trained deep neural network acoustic models on various datasets with a variety of architectures and loss functions. We show that BL-JUST can outperform the widely-used pre-training and fine-tuning strategy and some other popular semi-supervised techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。