arXiv:2410.08828cs.CLcs.SD2024-10中稿 · O-COCOSDA 2024被引 8

提升印尼语语音识别,测试多语言模型在多样语音下的表现

Enhancing Indonesian Automatic Speech Recognition: Evaluating Multilingual Models with Diverse Speech Variabilities

  • 用包含多种语音变体的印尼语数据集评估MMS和Whisper模型
  • Whisper微调后在不同语音条件下WER和CER最低,效果最佳
  • 说话风格差异对模型性能影响最大,适合关注语音多样性研究者

理想的语音识别模型应能准确转录不同语音特征的语音信号,如朗读与即兴说话、正式与非正式语境、干净与中等噪声环境。构建此类模型需大量具有多样化语音特征的训练数据。目前印尼语数据以朗读、正式、干净语音为主,缺乏其他语音变体。为推进印尼语自动语音识别(ASR),我们研究了最先进的语音识别模型Massively Multilingual Speech(MMS)和Whisper,同时构建了一个包含多样化语音变体的印尼语数据集以支持研究。我们进一步评估了这些模型在不同语音变体组上的预测能力。结果表明,经过微调的Whisper模型在各类数据集上表现最佳,显著降低了词错误率(WER)和字符错误率(CER)。此外,说话风格的差异对模型性能影响最为显著。

原文摘要 · Abstract (English)

An ideal speech recognition model has the capability to transcribe speech accurately under various characteristics of speech signals, such as speaking style (read and spontaneous), speech context (formal and informal), and background noise conditions (clean and moderate). Building such a model requires a significant amount of training data with diverse speech characteristics. Currently, Indonesian data is dominated by read, formal, and clean speech, leading to a scarcity of Indonesian data with other speech variabilities. To develop Indonesian automatic speech recognition (ASR), we present our research on state-of-the-art speech recognition models, namely Massively Multilingual Speech (MMS) and Whisper, as well as compiling a dataset comprising Indonesian speech with variabilities to facilitate our study. We further investigate the models' predictive ability to transcribe Indonesian speech data across different variability groups. The best results were achieved by the Whisper fine-tuned model across datasets with various characteristics, as indicated by the decrease in word error rate (WER) and character error rate (CER). Moreover, we found that speaking style variability affected model performance the most.

语音识别多语言模型印尼语语音变体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。