arXiv:2412.10857cs.SDcs.CV2024-12被引 2

用混合模型提升嘈杂环境下波斯数字的识别准确率

Robust Persian Digit Recognition in Noisy Environments Using Hybrid CNN-BiGRU Model

  • 结合残差卷积与双向GRU,用词单位实现免训练识别
  • 在噪声中达95.92%测试准确率,比LSTM模型高26.88%
  • 适合语音识别、低资源语言处理的研究者参考

人工智能在语音识别领域取得显著进展,但现有基于神经网络的方法在噪声环境中表现不佳。本研究针对孤立口语波斯数字(0至9)在噪声条件下的识别问题,特别是发音相似数字的区分难题,提出一种融合残差卷积神经网络与双向门控循环单元(BiGRU)的混合模型,采用词单位而非音素单位实现说话人无关识别。使用经过多种方式增强的FARSDIGIT1数据集,通过梅尔频率倒谱系数(MFCC)提取特征。实验结果表明,该模型在训练、验证和测试集上的准确率分别达到98.53%、96.10%和95.92%。在噪声条件下,相比基于音素的LSTM模型,识别准确率提升26.88%,优于使用梅尔尺度二维根倒谱系数(MTDRCC)与多层感知机(MLP)组合的方法(MTDRCC+MLP)7.61%。

原文摘要 · Abstract (English)

Artificial intelligence (AI) has significantly advanced speech recognition applications. However, many existing neural network-based methods struggle with noise, reducing accuracy in real-world environments. This study addresses isolated spoken Persian digit recognition (zero to nine) under noisy conditions, particularly for phonetically similar numbers. A hybrid model combining residual convolutional neural networks and bidirectional gated recurrent units (BiGRU) is proposed, utilizing word units instead of phoneme units for speaker-independent recognition. The FARSDIGIT1 dataset, augmented with various approaches, is processed using Mel-Frequency Cepstral Coefficients (MFCC) for feature extraction. Experimental results demonstrate the model's effectiveness, achieving 98.53%, 96.10%, and 95.92% accuracy on training, validation, and test sets, respectively. In noisy conditions, the proposed approach improves recognition by 26.88% over phoneme unit-based LSTM models and surpasses the Mel-scale Two Dimension Root Cepstrum Coefficients (MTDRCC) feature extraction technique along with MLP model (MTDRCC+MLP) by 7.61%.

语音识别深度学习噪声鲁棒波斯语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。