arXiv:2409.18428cs.CLcs.SD2024-09被引 2

用外部特征重排N-best结果,提升真实场景下的多语言语音识别准确率。

Improving Multilingual ASR in the Wild Using Simple N-best Re-ranking

  • 通过语言模型和文本识语模型重排N-best候选
  • 在FLEURS上降低3.3%和2.0%的词错误率
  • 适合实际应用中语言识别不准的多语言场景

多语言自动语音识别(ASR)模型通常在已知语种的设定下评估,但实际场景中语言标识常不准确。自动语音语言识别(SLID)模型存在误判,严重影响最终识别准确率。本文提出一种简单有效的N-best重排方法,利用语言模型和基于文本的语言识别模型等外部特征,改进多个主流声学模型的多语言ASR性能。在FLEURS数据集上,使用MMS和Whisper模型时,语音语言识别准确率分别提升8.7%和6.1%,词错误率分别降低3.3%和2.0%。

原文摘要 · Abstract (English)

Multilingual Automatic Speech Recognition (ASR) models are typically evaluated in a setting where the ground-truth language of the speech utterance is known, however, this is often not the case for most practical settings. Automatic Spoken Language Identification (SLID) models are not perfect and misclassifications have a substantial impact on the final ASR accuracy. In this paper, we present a simple and effective N-best re-ranking approach to improve multilingual ASR accuracy for several prominent acoustic models by employing external features such as language models and text-based language identification models. Our results on FLEURS using the MMS and Whisper models show spoken language identification accuracy improvements of 8.7% and 6.1%, respectively and word error rates which are 3.3% and 2.0% lower on these benchmarks.

多语言识别语音识别N-best重排

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。