跨场景语音提取新方法,让模型一次训练通用于多种噪声环境。
TripleC Learning and Lightweight Speech Enhancement for Multi-Condition Target Speech Extraction
- 引入跨条件一致性学习,统一训练多类语音混合场景。
- 在Libri2Mix三场景测试中超越专用模型,性能更优。
- 轻量级设计适合真实复杂环境部署,尤其适合资源受限设备。
在先前工作中,我们提出了轻量级语音增强引导的目标语音提取方法(LGTSE),并在多说话人加噪声场景中验证了其有效性。然而,现实应用常涉及更复杂的条件组合,如单说话人加噪声或双说话人无噪声。为此,我们扩展LGTSE,引入一种跨条件一致性学习策略,称为TripleC学习。该策略首先在多说话人加噪声条件下验证,随后评估其在多样化场景下的泛化能力。基于LGTSE中轻量级前端去噪器可灵活处理含噪与纯净混合信号且对未见条件具有良好泛化性的特性,我们进一步提出并行通用训练方案,将包含多种场景的批次数据统一用于同一目标说话人的训练。通过强制不同条件下提取结果的一致性,简单任务可辅助困难任务,充分挖掘多样训练数据潜力,构建鲁棒的通用模型。在Libri2Mix三条件任务上的实验表明,所提LGTSE结合TripleC学习的方法在性能上优于各类条件专用模型,展现出在真实语音应用中通用部署的强大潜力。
原文摘要 · Abstract (English)
In our recent work, we proposed Lightweight Speech Enhancement Guided Target Speech Extraction (LGTSE) and demonstrated its effectiveness in multi-speaker-plus-noise scenarios. However, real-world applications often involve more diverse and complex conditions, such as one-speaker-plus-noise or two-speaker-without-noise. To address this challenge, we extend LGTSE with a Cross-Condition Consistency learning strategy, termed TripleC Learning. This strategy is first validated under multi-speaker-plus-noise condition and then evaluated for its generalization across diverse scenarios. Moreover, building upon the lightweight front-end denoiser in LGTSE, which can flexibly process both noisy and clean mixtures and shows strong generalization to unseen conditions, we integrate TripleC learning with a proposed parallel universal training scheme that organizes batches containing multiple scenarios for the same target speaker. By enforcing consistent extraction across different conditions, easier cases can assist harder ones, thereby fully exploiting diverse training data and fostering a robust universal model. Experimental results on the Libri2Mix three-condition tasks demonstrate that the proposed LGTSE with TripleC learning achieves superior performance over condition-specific models, highlighting its strong potential for universal deployment in real-world speech applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。