arXiv:2505.01632eess.AScs.AI2025-05被引 5

用迁移学习的残差网络提升语音识别在嘈杂环境下的准确率

Transfer Learning-Based Deep Residual Learning for Speech Recognition in Clean and Noisy Environments

  • 基于残差网络的迁移学习框架,增强噪声下语音特征提取能力
  • 在干净环境下识别率达98.94%,噪声环境下达91.21%
  • 适合需要鲁棒语音识别的工业场景应用

针对非平稳环境噪声对自动语音识别(ASR)的负面影响,本文提出一种融合稳健前端的新型神经网络框架,适用于清晰与嘈杂环境。基于Aurora-2语音数据集,采用梅尔频率声学特征,结合基于残差网络(ResNet)的迁移学习方法进行评估。实验表明,该方法在识别准确率上显著优于卷积神经网络(CNN)和长短期记忆网络(LSTM)。在干净环境中达到98.94%的准确率,在噪声环境下实现91.21%的准确率。

原文摘要 · Abstract (English)

Addressing the detrimental impact of non-stationary environmental noise on automatic speech recognition (ASR) has been a persistent and significant research focus. Despite advancements, this challenge continues to be a major concern. Recently, data-driven supervised approaches, such as deep neural networks, have emerged as promising alternatives to traditional unsupervised methods. With extensive training, these approaches have the potential to overcome the challenges posed by diverse real-life acoustic environments. In this light, this paper introduces a novel neural framework that incorporates a robust frontend into ASR systems in both clean and noisy environments. Utilizing the Aurora-2 speech database, the authors evaluate the effectiveness of an acoustic feature set for Mel-frequency, employing the approach of transfer learning based on Residual neural network (ResNet). The experimental results demonstrate a significant improvement in recognition accuracy compared to convolutional neural networks (CNN) and long short-term memory (LSTM) networks. They achieved accuracies of 98.94% in clean and 91.21% in noisy mode.

语音识别迁移学习残差网络噪声鲁棒

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。