arXiv:2410.22903eess.AScs.SD2024-10

用合成语音提升波兰语语音识别准确率

Augmenting Polish Automatic Speech Recognition System With Synthetic Data

  • 基于Voicebox生成合成语音数据增强模型训练
  • 合成数据使Conformer和Whisper模型性能显著提升
  • 适用于语音识别数据稀缺场景的改进方案

本文介绍为参加Poleval 2024任务3:波兰语自动语音识别挑战赛而开发的系统。我们描述了基于Voicebox的语音合成流程,并利用其生成的合成语音数据对Conformer和Whisper语音识别模型进行增强。实验表明,在训练中加入合成语音可显著提升模型表现。文中还展示了我们的模型在竞赛中取得的最终结果。

原文摘要 · Abstract (English)

This paper presents a system developed for submission to Poleval 2024, Task 3: Polish Automatic Speech Recognition Challenge. We describe Voicebox-based speech synthesis pipeline and utilize it to augment Conformer and Whisper speech recognition models with synthetic data. We show that addition of synthetic speech to training improves achieved results significantly. We also present final results achieved by our models in the competition.

语音识别合成数据波兰语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。