arXiv:2409.01545cs.SDcs.AI2024-09中稿 · IEEE SLT 2024被引 3

用少量目标域噪音数据生成逼真语音,提升跨域语音增强效果

Effective Noise-aware Data Simulation for Domain-adaptive Speech Enhancement Leveraging Dynamic Stochastic Perturbation

  • 通过噪声编码器提取目标域噪声特征,指导生成器合成匹配的语音
  • 引入动态随机扰动机制,在推理时注入可控噪声扰动以增强泛化能力
  • 仅需少量目标域数据即可实现良好跨域性能,适合真实场景部署

跨域语音增强常因未知目标域中缺乏噪声和背景信息而面临严峻挑战,导致训练与测试条件不匹配。本文提出一种新型数据模拟方法,仅需有限的目标域含噪语音数据,结合噪声提取技术与生成对抗网络(GAN)来缓解此问题。核心思想是利用噪声编码器从目标域数据中提取噪声嵌入,引导生成器合成声学特性适配目标域的语音,同时真实保留输入干净语音的语音内容。此外,本文引入动态随机扰动概念,在推理阶段对噪声嵌入施加可控扰动,使模型能有效泛化至未见噪声环境。在VoiceBank-DEMAND基准数据集上的实验表明,所提方法在跨域语音增强任务上优于现有强基线方法。

原文摘要 · Abstract (English)

Cross-domain speech enhancement (SE) is often faced with severe challenges due to the scarcity of noise and background information in an unseen target domain, leading to a mismatch between training and test conditions. This study puts forward a novel data simulation method to address this issue, leveraging noise-extractive techniques and generative adversarial networks (GANs) with only limited target noisy speech data. Notably, our method employs a noise encoder to extract noise embeddings from target-domain data. These embeddings aptly guide the generator to synthesize utterances acoustically fitted to the target domain while authentically preserving the phonetic content of the input clean speech. Furthermore, we introduce the notion of dynamic stochastic perturbation, which can inject controlled perturbations into the noise embeddings during inference, thereby enabling the model to generalize well to unseen noise conditions. Experiments on the VoiceBank-DEMAND benchmark dataset demonstrate that our domain-adaptive SE method outperforms an existing strong baseline based on data simulation.

语音增强域适应数据模拟生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。