arXiv:2505.12288eess.AScs.SD2025-05被引 1

统一架构让语音增强与个性化增强共用模型,提升效果且部署更简单。

Unified Architecture and Unsupervised Speech Disentanglement for Speaker Embedding-Free Enrollment in Personalized Speech Enhancement

  • 构建统一框架,同时处理常规增强与个性化增强任务。
  • 通过无监督解耦,分离说话人身份信息,提升抗情绪干扰能力。
  • 探索长短录音配对策略,发现随机时长录音效果更优。

传统语音增强(SE)旨在抑制噪声以改善语音感知和可懂度,无需参考录音;而个性化语音增强(PSE)则利用录音提取目标说话人语音,解决鸡尾酒会问题。两者虽目标不同但模型架构相似,均包含处理录音的分支。为此,本文提出USEF-PNet和DSEF-PNet两个新模型,基于此前的SEF-PNet框架。USEF-PNet采用统一架构整合SE与PSE,提升性能并简化部署。DSEF-PNet引入无监督语音解耦机制:将混合语音与两段不同录音配对,强制提取的目标语音保持一致,从而有效分离高质量说话人身份特征,减少情绪、内容等干扰,增强PSE鲁棒性。此外,研究了长-短录音配对(LSEP)策略,考察录音时长在训练与评估中的影响。在Libri2Mix和VoiceBank DEMAND数据集上的实验表明,所提模型显著提升性能,随机时长录音表现略优。

原文摘要 · Abstract (English)

Conventional speech enhancement (SE) aims to improve speech perception and intelligibility by suppressing noise without requiring enrollment speech as reference, whereas personalized SE (PSE) addresses the cocktail party problem by extracting a target speaker's speech using enrollment speech. While these two tasks tackle different yet complementary challenges in speech signal processing, they often share similar model architectures, with PSE incorporating an additional branch to process enrollment speech. This suggests developing a unified model capable of efficiently handling both SE and PSE tasks, thereby simplifying deployment while maintaining high performance. However, PSE performance is sensitive to variations in enrollment speech, like emotional tone, which limits robustness in real-world applications. To address these challenges, we propose two novel models, USEF-PNet and DSEF-PNet, both extending our previous SEF-PNet framework. USEF-PNet introduces a unified architecture for processing enrollment speech, integrating SE and PSE into a single framework to enhance performance and streamline deployment. Meanwhile, DSEF-PNet incorporates an unsupervised speech disentanglement approach by pairing a mixture speech with two different enrollment utterances and enforcing consistency in the extracted target speech. This strategy effectively isolates high-quality speaker identity information from enrollment speech, reducing interference from factors such as emotion and content, thereby improving PSE robustness. Additionally, we explore a long-short enrollment pairing (LSEP) strategy to examine the impact of enrollment speech duration during both training and evaluation. Extensive experiments on the Libri2Mix and VoiceBank DEMAND demonstrate that our proposed USEF-PNet, DSEF-PNet all achieve substantial performance improvements, with random enrollment duration performing slightly better.

语音增强个性化语音无监督学习统一模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。