arXiv:2602.04307eess.AScs.CL2026-02中稿 · IEEE Transactions …

统一框架解决语音识别与增强中的噪声和信道失配问题

Universal Robust Speech Adaptation for Cross-Domain Speech Recognition and Enhancement

  • 用双编码器分别学习噪声与信道特征,生成匹配目标域的语音
  • 在复杂混合干扰下,语音识别错误率降低16.16%,语音增强感知指标提升15.58%
  • 动态随机扰动技术增强模型对未见域的鲁棒性,适合实际部署场景

用于自动语音识别(ASR)和语音增强(SE)的预训练模型在匹配噪声和信道条件下表现优异,但在遭遇领域偏移时性能显著下降,尤其面对未见过的噪声与信道畸变。为此,本文提出URSA-GAN,一种统一且领域感知的生成式框架,旨在缓解噪声与信道条件不匹配问题。URSA-GAN采用双嵌入架构,包含噪声编码器与信道编码器,均通过有限领域内数据预训练以捕捉相关领域表征。这些嵌入条件化一个基于GAN的语音生成器,实现与目标域声学特性一致的同时保留语音内容。为进一步提升泛化能力,提出动态随机扰动,一种新型正则化技术,在生成过程中向嵌入引入可控变异,增强对未见领域的鲁棒性。实验表明,URSA-GAN在多种噪声与不匹配信道场景下有效降低ASR字符错误率,并提升SE感知指标。在同时存在信道与噪声退化的复合测试条件下,其表现出强泛化能力,相对提升达16.16%(ASR)与15.58%(SE)。

原文摘要 · Abstract (English)

Pre-trained models for automatic speech recognition (ASR) and speech enhancement (SE) have exhibited remarkable capabilities under matched noise and channel conditions. However, these models often suffer from severe performance degradation when confronted with domain shifts, particularly in the presence of unseen noise and channel distortions. In view of this, we in this paper present URSA-GAN, a unified and domain-aware generative framework specifically designed to mitigate mismatches in both noise and channel conditions. URSA-GAN leverages a dual-embedding architecture that consists of a noise encoder and a channel encoder, each pre-trained with limited in-domain data to capture domain-relevant representations. These embeddings condition a GAN-based speech generator, facilitating the synthesis of speech that is acoustically aligned with the target domain while preserving phonetic content. To enhance generalization further, we propose dynamic stochastic perturbation, a novel regularization technique that introduces controlled variability into the embeddings during generation, promoting robustness to unseen domains. Empirical results demonstrate that URSA-GAN effectively reduces character error rates in ASR and improves perceptual metrics in SE across diverse noisy and mismatched channel scenarios. Notably, evaluations on compound test conditions with both channel and noise degradations confirm the generalization ability of URSA-GAN, yielding relative improvements of 16.16% in ASR performance and 15.58% in SE metrics.

语音识别语音增强域适应生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。