arXiv:2510.25150cs.CL2025-10被引 2

分离语音语义与噪声,提升嘈杂环境下的语音识别准确率

Explainable Disentanglement on Discrete Speech Representations for Noise-Robust ASR

  • 将语音嵌入分解为语义代码本令牌和可解释的噪声向量
  • 在VBDemand数据集上错误率降低82%,优于基线35%
  • 适合需要鲁棒语音识别与可解释性的实际应用场景

离散语音表示因其可解释性与大语言模型的兼容性,在语音建模中日益受到关注,但在噪声或真实环境中的表现仍不理想。基于将Whisper嵌入量化为语音-单元建模的方法,我们提出在潜在空间中分离语音语义与背景噪声。所提端到端模型将纯净语音表示为码本令牌,同时提取可解释的噪声向量作为量化残差,并通过轻量级分类器进行监督。实验表明,该方法显著提升了干净与嘈杂语音及文本之间的对齐度,生成的语音令牌具有高度抗噪性,有效改善了自动语音识别性能。在保持Whisper冻结的前提下,相比Whisper在VBDemand测试集上错误率降低82%,较基线方法提升35%。进一步分析显示,学习到的令牌空间在已见和未见声学条件下均具有良好泛化能力。

原文摘要 · Abstract (English)

Discrete audio representations are gaining traction in speech modeling due to their interpretability and compatibility with large language models, but are not always optimized for noisy or real-world environments. Building on existing works that quantize Whisper embeddings for speech-to-unit modeling, we propose disentangling semantic speech content from background noise in the latent space. Our end-to-end model separates clean speech in the form of codebook tokens, while extracting interpretable noise vectors as quantization residue which are supervised via a lightweight classifier. We show that our approach improves alignment between clean/noisy speech and text, producing speech tokens that display a high degree of noiseinvariance, and improves ASR performance. Keeping Whisper frozen, we show an 82% reduction in error rate compared to Whisper, and 35% improvement over baseline methods on the VBDemand test set. Further analyses show that the learned token space generalizes well to both seen and unseen acoustic conditions.

语音识别降噪可解释性离散表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。