arXiv:2505.04237eess.AScs.SD2025-05被引 1

用薛定谔桥模型增强语音,显著提升嘈杂环境下的识别准确率。

Robust Speech Recognition with Schrödinger Bridge-Based Speech Enhancement

  • 基于薛定谔桥生成语音增强,替代传统扩散模型。
  • 相比原始语音,词错误率降低约40%;比同类预测模型降8%。
  • 适合需要高鲁棒性的语音识别系统部署。

本文研究将生成式语音增强应用于提升语音识别模型在噪声和混响环境下的鲁棒性。采用一种基于薛定谔桥的新型语音增强模型,该模型相较于扩散模型表现更优。我们分析了模型规模与不同采样方法对语音识别性能的影响,并与预测型及扩散基基线模型进行了比较,考察了在使用不同预训练语音识别模型时的表现。所提方法显著降低了词错误率:相比未处理语音信号减少约40%,相比同规模预测型方法减少约8%。

原文摘要 · Abstract (English)

In this work, we investigate application of generative speech enhancement to improve the robustness of ASR models in noisy and reverberant conditions. We employ a recently-proposed speech enhancement model based on Schrödinger bridge, which has been shown to perform well compared to diffusion-based approaches. We analyze the impact of model scaling and different sampling methods on the ASR performance. Furthermore, we compare the considered model with predictive and diffusion-based baselines and analyze the speech recognition performance when using different pre-trained ASR models. The proposed approach significantly reduces the word error rate, reducing it by approximately 40% relative to the unprocessed speech signals and by approximately 8% relative to a similarly sized predictive approach.

语音增强薛定谔桥语音识别鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。