arXiv:2512.08973cs.SDcs.AI2025-12

将噪声检测嵌入语音识别模型,提升嘈杂环境下的识别准确率。

Enhancing Automatic Speech Recognition Through Integrated Noise Detection Architecture

  • 在wav2vec2框架中加入并行噪声识别模块,同步处理语音与噪声。
  • 词错误率和字符错误率显著降低,噪声检测准确率达92.3%。
  • 适合需要高鲁棒性的实时语音识别系统开发者使用。

本研究提出一种新方法,通过将噪声检测能力直接集成到语音识别架构中,提升自动语音识别系统的性能。基于wav2vec2框架,该方法引入一个专用的噪声识别模块,与语音转录过程并行运行。在公开的语音与环境音频数据集上进行实验验证,结果表明,该增强系统在词错误率、字符错误率及噪声检测准确率方面均优于传统架构。实验显示,联合优化语音转录与噪声分类目标,能显著提升复杂声学环境下语音识别的可靠性。

原文摘要 · Abstract (English)

This research presents a novel approach to enhancing automatic speech recognition systems by integrating noise detection capabilities directly into the recognition architecture. Building upon the wav2vec2 framework, the proposed method incorporates a dedicated noise identification module that operates concurrently with speech transcription. Experimental validation using publicly available speech and environmental audio datasets demonstrates substantial improvements in transcription quality and noise discrimination. The enhanced system achieves superior performance in word error rate, character error rate, and noise detection accuracy compared to conventional architectures. Results indicate that joint optimization of transcription and noise classification objectives yields more reliable speech recognition in challenging acoustic conditions.

语音识别噪声抑制深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。