arXiv:2507.21331cs.CLcs.SD2025-07被引 3

用深度学习提升绍纳语语音识别准确率,解决数据少、声调复杂难题。

A Deep Learning Automatic Speech Recognition Model for Shona Language

  • 混合CNN+LSTM架构结合注意力机制,适应绍纳语声调特点。
  • 通过数据增强与迁移学习,在有限数据下实现29%词错误率。
  • 为低资源语言语音识别提供可复用的技术方案,适合语言技术研究者。

本研究开发了一种基于深度学习的绍纳语自动语音识别系统,针对该语言因数据稀缺、标注数据不足及独特的声调与语法复杂性带来的挑战。研究首先验证了深度学习在绍纳语语音识别中的可行性;其次探讨了设计与实现深度神经网络架构的具体难点,并提出相应缓解策略;最后对比了深度学习模型与传统统计模型的识别性能。所构建的语音识别系统采用卷积神经网络进行声学建模,长短期记忆网络进行语言建模,结合注意力机制以应对声调特性。为克服数据短缺问题,引入数据增强和迁移学习技术。实验结果显示,系统在测试集上达到29%的词错误率(Word Error Rate)、12%的音素错误率(Phoneme Error Rate)以及74%的整体准确率。结果表明,深度学习可显著提升低资源语言如绍纳语的语音识别性能。本研究推动了低资源语言语音识别技术的发展,有助于提升全球绍纳语使用者的交流与可及性。

原文摘要 · Abstract (English)

This study presented the development of a deep learning-based Automatic Speech Recognition system for Shona, a low-resource language characterized by unique tonal and grammatical complexities. The research aimed to address the challenges posed by limited training data, lack of labelled data, and the intricate tonal nuances present in Shona speech, with the objective of achieving significant improvements in recognition accuracy compared to traditional statistical models. The research first explored the feasibility of using deep learning to develop an accurate ASR system for Shona. Second, it investigated the specific challenges involved in designing and implementing deep learning architectures for Shona speech recognition and proposed strategies to mitigate these challenges. Lastly, it compared the performance of the deep learning-based model with existing statistical models in terms of accuracy. The developed ASR system utilized a hybrid architecture consisting of a Convolutional Neural Network for acoustic modelling and a Long Short-Term Memory network for language modelling. To overcome the scarcity of data, data augmentation techniques and transfer learning were employed. Attention mechanisms were also incorporated to accommodate the tonal nature of Shona speech. The resulting ASR system achieved impressive results, with a Word Error Rate of 29%, Phoneme Error Rate of 12%, and an overall accuracy of 74%. These metrics indicated the potential of deep learning to enhance ASR accuracy for under-resourced languages like Shona. This study contributed to the advancement of ASR technology for under-resourced languages like Shona, ultimately fostering improved accessibility and communication for Shona speakers worldwide.

语音识别深度学习低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。