用7层神经网络实现93.4%准确率的孟加拉语孤立词语音识别
Deep Feed-Forward Neural Network for Bangla Isolated Speech Recognition
- 采用MFCC特征与7层全连接网络进行分类
- 在自建数据集上达到93.42%识别准确率
- 为低资源语言语音识别提供可行方案
作为重要的人机交互工具,相较于英语,针对孟加拉语的语音识别研究仍十分有限。为此,本文构建了一个包含孤立孟加拉语和英语词汇的语料库,并基于该数据集实现并分析了无发言人依赖的孤立词语音识别系统。提出一种结合梅尔频率倒谱系数(MFCC)与7层深度前馈全连接神经网络(DFFNN)的识别方法。实验结果表明,该方法在考虑类别数量和数据集规模的情况下,达到了93.42%的识别准确率,优于此前多数孟加拉语语音识别研究。
原文摘要 · Abstract (English)
As the most important human-machine interfacing tool, an insignificant amount of work has been carried out on Bangla Speech Recognition compared to the English language. Motivated by this, in this work, the performance of speaker-independent isolated speech recognition systems has been implemented and analyzed using a dataset that is created containing both isolated Bangla and English spoken words. An approach using the Mel Frequency Cepstral Coefficient (MFCC) and Deep Feed-Forward Fully Connected Neural Network (DFFNN) of 7 layers as a classifier is proposed in this work to recognize isolated spoken words. This work shows 93.42% recognition accuracy which is better compared to most of the works done previously on Bangla speech recognition considering the number of classes and dataset size.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。