arXiv:2508.19583eess.AS2025-08被引 1

轻量级语音增强模型提升嘈杂多说话人场景下的目标语音提取效果

Lightweight speech enhancement guided target speech extraction in noisy multi-speaker scenarios

  • 用轻量级模型GTCRN先降噪,再引导目标语音提取
  • 在Libri2Mix上实现SISDR提升0.89 dB,PESQ提升0.16
  • 适合需要低延迟、高鲁棒性的实时语音分离场景

目标语音提取(TSE)在单说话人加噪声和双说话人混合等简单场景中表现良好,但在嘈杂的多说话人场景下性能仍不理想。为此,我们引入轻量级语音增强模型GTCRN,以更好引导TSE在噪声环境中的表现。在先前无需说话人嵌入/编码器的SEF-PNet框架基础上,提出两种扩展:LGTSE和D-LGTSE。LGTSE通过在上下文交互前对输入噪声语音进行降噪,实现与说话人无关的注册引导,从而降低噪声干扰。D-LGTSE进一步通过在训练时使用降噪后的语音作为额外噪声输入,扩大噪声条件的动态范围,使模型能直接学习受损信号。此外,采用两阶段训练策略:先进行GTCRN增强引导的预训练,再联合微调,以充分挖掘模型潜力。在Libri2Mix数据集上的实验表明,该方法在SISDR上提升0.89 dB,PESQ提升0.16,STOI提升1.97%,验证了方法的有效性。

原文摘要 · Abstract (English)

Target speech extraction (TSE) has achieved strong performance in relatively simple conditions such as one-speaker-plus-noise and two-speaker mixtures, but its performance remains unsatisfactory in noisy multi-speaker scenarios. To address this issue, we introduce a lightweight speech enhancement model, GTCRN, to better guide TSE in noisy environments. Building on our competitive previous speaker embedding/encoder-free framework SEF-PNet, we propose two extensions: LGTSE and D-LGTSE. LGTSE incorporates noise-agnostic enrollment guidance by denoising the input noisy speech before context interaction with enrollment speech, thereby reducing noise interference. D-LGTSE further improves system robustness against speech distortion by leveraging denoised speech as an additional noisy input during training, expanding the dynamic range of noisy conditions and enabling the model to directly learn from distorted signals. Furthermore, we propose a two-stage training strategy, first with GTCRN enhancement-guided pre-training and then joint fine-tuning, to fully exploit model potential.Experiments on the Libri2Mix dataset demonstrate significant improvements of 0.89 dB in SISDR, 0.16 in PESQ, and 1.97% in STOI, validating the effectiveness of our approach.

语音增强多说话人轻量模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。