通过融合混合语音与注册语音的上下文互动,提升目标说话人提取精度。
DualStream Contextual Fusion Network: Efficient Target Speaker Extraction by Leveraging Mixture and Enrollment Interactions
- 设计双流结构捕捉混合语音与注册语音的时空交互
- 在基准数据集上实现21.6 dB的SI-SDRi提升
- 误提取率低至0.4%,适合实际部署场景
目标说话人提取旨在利用注册信息从多人混响环境中分离出目标语音。现有方法多依赖注册语音获得的说话人嵌入,可能忽略上下文信息及混合语音与注册语音间的内部交互。本文提出一种新型双流上下文融合网络(DCF-Net),在时频域中建模。具体地,引入双流融合模块(DSFB),捕获上下文化注册语音与混合语音表示在空间和通道维度上的交互,生成丰富且一致的表征以指导提取网络。实验表明,DCF-Net优于当前最先进方法,在基准数据集上实现21.6 dB的尺度不变信号失真比改进(SI-SDRi),在噪声与混响场景下均表现鲁棒有效。此外,模型错误提取率(目标混淆问题)降至0.4%,展现出在实际应用中的潜力。
原文摘要 · Abstract (English)
Target speaker extraction focuses on extracting a target speech signal from an environment with multiple speakers by leveraging an enrollment. Existing methods predominantly rely on speaker embeddings obtained from the enrollment, potentially disregarding the contextual information and the internal interactions between the mixture and enrollment. In this paper, we propose a novel DualStream Contextual Fusion Network (DCF-Net) in the time-frequency (T-F) domain. Specifically, DualStream Fusion Block (DSFB) is introduced to obtain contextual information and capture the interactions between contextualized enrollment and mixture representation across both spatial and channel dimensions, and then rich and consistent representations are utilized to guide the extraction network for better extraction. Experimental results demonstrate that DCF-Net outperforms state-of-the-art (SOTA) methods, achieving a scale-invariant signal-to-distortion ratio improvement (SI-SDRi) of 21.6 dB on the benchmark dataset, and exhibits its robustness and effectiveness in both noise and reverberation scenarios. In addition, the wrong extraction results of our model, called target confusion problem, reduce to 0.4%, which highlights the potential of DCF-Net for practical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。