arXiv:2502.17911cs.SDcs.AI2025-02被引 1

融合BGRU与Transformer提升语音降噪效果,更清晰。

Enhancing Speech Quality through the Integration of BGRU and Transformer Architectures

  • 用双向门控循环单元与变换器联合建模语音时序特征
  • 在噪声环境下显著提升语音质量,优于单一模型
  • 适合需要高鲁棒性的语音增强应用场景

语音增强在噪声环境中对提升语音质量至关重要。本文研究了将双向门控循环单元(BGRU)与Transformer模型结合用于语音增强任务的可行性。通过全面的实验评估,结果表明该混合架构在性能上优于传统方法及独立模型。BGRU-Transformer框架在捕捉时序依赖性和学习复杂信号模式方面表现优异,有效提升了降噪能力与语音质量。实验结果显示,相比现有方法有显著性能提升,凸显该集成模型在实际应用中的潜力。BGRU与Transformer的无缝融合不仅增强了系统鲁棒性,也为先进语音处理技术的发展提供了新路径。本研究推动了语音增强技术的发展,为优化模型架构、探索多样化应用场景及深化噪声环境下的语音处理研究奠定了坚实基础。

原文摘要 · Abstract (English)

Speech enhancement plays an essential role in improving the quality of speech signals in noisy environments. This paper investigates the efficacy of integrating Bidirectional Gated Recurrent Units (BGRU) and Transformer models for speech enhancement tasks. Through a comprehensive experimental evaluation, our study demonstrates the superiority of this hybrid architecture over traditional methods and standalone models. The combined BGRU-Transformer framework excels in capturing temporal dependencies and learning complex signal patterns, leading to enhanced noise reduction and improved speech quality. Results show significant performance gains compared to existing approaches, highlighting the potential of this integrated model in real-world applications. The seamless integration of BGRU and Transformer architectures not only enhances system robustness but also opens the road for advanced speech processing techniques. This research contributes to the ongoing efforts in speech enhancement technology and sets a solid foundation for future investigations into optimizing model architectures, exploring many application scenarios, and advancing the field of speech processing in noisy environments.

语音增强深度学习序列建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。