无需说话人编码器,通过上下文融合提升语音增强效果
SEF-PNet: Speaker Encoder-Free Personalized Speech Enhancement with Local and Global Contexts Aggregation
- 用动态交互与上下文聚合替代传统说话人编码器
- 在Libri2Mix上达到当前最优的个性化语音增强性能
- 适合追求轻量级、高精度语音增强的研究者
个性化语音增强(PSE)方法通常依赖预训练的说话人验证模型或自设计的说话人编码器来提取目标说话人特征,以指导模型分离出期望语音。然而,这些方法存在模型复杂度高、未充分利用注册说话人信息等问题,限制了性能潜力。为此,本文提出一种新型无说话人编码器的个性化语音增强网络——SEF-PNet,充分挖掘注册语音和噪声混合信号中的信息。SEF-PNet引入两项关键创新:交互式说话人适配(ISA)与局部-全局上下文聚合(LCA)。ISA动态调节注册信号与噪声信号间的交互,增强说话人适应能力;LCA在PSE编码器中采用先进通道注意力机制,有效整合局部与全局上下文信息,提升特征学习能力。在Libri2Mix数据集上的实验表明,SEF-PNet显著优于基线模型,达到当前最先进的PSE性能。
原文摘要 · Abstract (English)
Personalized speech enhancement (PSE) methods typically rely on pre-trained speaker verification models or self-designed speaker encoders to extract target speaker clues, guiding the PSE model in isolating the desired speech. However, these approaches suffer from significant model complexity and often underutilize enrollment speaker information, limiting the potential performance of the PSE model. To address these limitations, we propose a novel Speaker Encoder-Free PSE network, termed SEF-PNet, which fully exploits the information present in both the enrollment speech and noisy mixtures. SEF-PNet incorporates two key innovations: Interactive Speaker Adaptation (ISA) and Local-Global Context Aggregation (LCA). ISA dynamically modulates the interactions between enrollment and noisy signals to enhance the speaker adaptation, while LCA employs advanced channel attention within the PSE encoder to effectively integrate local and global contextual information, thus improving feature learning. Experiments on the Libri2Mix dataset demonstrate that SEF-PNet significantly outperforms baseline models, achieving state-of-the-art PSE performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。