arXiv:2604.11256eess.AS2026-04

同时更新教师模型组与学生模型,提升语音识别跨域适应性能。

Teaching the Teachers: Boosting unsupervised domain adaptation in speech recognition by ensemble update

  • 联合更新教师模型组与学生模型,避免分步训练
  • 在SwitchBoard数据集上降低4.6%词错误率
  • 适合需要高效跨域适配的语音系统开发者

语音识别系统在未参与训练的数据领域上表现不佳。为解决此问题,无监督域适应方法通过集成与多阶段师生训练降低词错误率(WER)。尽管有所改进,其性能仍远低于有监督域内训练。本文提出一种更高效策略:同时更新教师模型集合与单个学生模型,无需分步训练。联合更新提升了学生模型的WER,反过来也促进了教师模型的逐步优化。实验使用三个带标签源数据集(AMI、WSJ、LS360)和一个无标签目标域(SwitchBoard),结果表明该方法在SwitchBoard eval00测试集上将WER降低4.6%,优于多阶段与迭代训练方法。

原文摘要 · Abstract (English)

Speech recognition systems often struggle with data domains that have not been included in the training. To address this, unsupervised domain adaptation has been explored with ensemble and multi-stage teacher-student training methods reducing the word error rate. Despite improvements, the error rate remains much higher than that achieved with supervised in-domain training. This work proposes a more efficient strategy by simultaneously updating the ensemble of teacher models along with the single student model eliminating the need for sequential models training. The joint update improves the word error rate of the student model, benefiting the progressively enhanced teacher models. Experiments are conducted with three labelled source datasets, namely AMI, WSJ, LS360, and one unlabeled target domain i.e. SwitchBoard. The results show that the proposed method improves the WER by 4.6% on the Switchboard eval00 test set, thus outperforming multi-stage and iterative training methods.

语音识别域适应师生训练集成学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。