arXiv:2511.03361eess.AScs.AI2025-11被引 2

基于FastConformer的开源罗马尼亚语语音识别系统,准确率提升27%。

Open Source State-Of-the-Art Solution for Romanian Speech Recognition

  • 首次在罗马尼亚语中使用FastConformer架构,结合弱监督数据训练。
  • 在读写、口语和领域特定语音上均达到最优,相对错误率降低27%。
  • 支持低延迟部署,适合科研与实际应用,解码效率高。

本文提出一种基于NVIDIA FastConformer架构的罗马尼亚语自动语音识别(ASR)系统,首次在该语言中探索此架构的应用。模型在超过2,600小时的弱监督转录语料上进行训练,采用融合连接时序分类(CTC)与词元持续时间转换器(TDT)的混合解码器,评估了贪婪解码、ALSD及带六元语法语言模型的CTC束搜索等策略。该系统在所有罗马尼亚语评测基准上均达到最先进性能,涵盖朗读、即兴和领域特定语音,相对此前最佳系统实现最高27%的相对词错误率(WER)下降。除准确性提升外,方法还展现出良好的解码效率,适用于低延迟语音识别场景的研究与部署。

原文摘要 · Abstract (English)

In this work, we present a new state-of-the-art Romanian Automatic Speech Recognition (ASR) system based on NVIDIA's FastConformer architecture--explored here for the first time in the context of Romanian. We train our model on a large corpus of, mostly, weakly supervised transcriptions, totaling over 2,600 hours of speech. Leveraging a hybrid decoder with both Connectionist Temporal Classification (CTC) and Token-Duration Transducer (TDT) branches, we evaluate a range of decoding strategies including greedy, ALSD, and CTC beam search with a 6-gram token-level language model. Our system achieves state-of-the-art performance across all Romanian evaluation benchmarks, including read, spontaneous, and domain-specific speech, with up to 27% relative WER reduction compared to previous best-performing systems. In addition to improved transcription accuracy, our approach demonstrates practical decoding efficiency, making it suitable for both research and deployment in low-latency ASR applications.

语音识别罗马尼亚语FastConformer弱监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。