arXiv:2502.16611cs.SDcs.AI2025-02NeurIPS被引 3

用噪声录音对比正负片段,实现从混响语音中提取目标说话人。

Target Speaker Extraction through Comparing Noisy Positive and Negative Audio Enrollments

  • 通过对比目标说话人发声与静默段落,从噪声录音中提取身份特征
  • 在双人混响语音中提升2.1 dB的SI-SNRi性能,优于已有方法
  • 适合实际场景中无干净录音时的目标说话人分离任务

目标说话人提取旨在从包含多个说话人的音频混合中分离出特定说话人的语音。以往方法依赖干净音频作为条件输入以提供目标身份信息,但此类清晰录音往往难以获取。例如,在嘈杂聚会中无法轻易获得陌生人的干净语音样本。现有研究较少探索从含干扰重叠语音的噪声录音中提取目标特征。本文提出一种新录入策略:通过比较目标说话人发声段(正样本)与静默段(负样本)来编码其身份信息。实验表明,所提模型架构在单声道语音混合中提取目标说话人时,比先前方法提升超过2.1 dB的SI-SNRi;此外,两阶段训练策略加速收敛,使达到3 dB SNR所需的优化步数减少60%。整体上,该方法在基于噪声录音的目标说话人提取任务中达到当前最优性能。代码已公开于https://github.com/xu-shitong/TSE-through-Positive-Negative-Enroll。

原文摘要 · Abstract (English)

Target speaker extraction focuses on isolating a specific speaker's voice from an audio mixture containing multiple speakers. To provide information about the target speaker's identity, prior works have utilized clean audio samples as conditioning inputs. However, such clean audio examples are not always readily available. For instance, obtaining a clean recording of a stranger's voice at a cocktail party without leaving the noisy environment is generally infeasible. Limited prior research has explored extracting the target speaker's characteristics from noisy enrollments, which may contain overlapping speech from interfering speakers. In this work, we explore a novel enrollment strategy that encodes target speaker information from the noisy enrollment by comparing segments where the target speaker is talking (Positive Enrollments) with segments where the target speaker is silent (Negative Enrollments). Experiments show the effectiveness of our model architecture, which achieves over 2.1 dB higher SI-SNRi compared to prior works in extracting the monaural speech from the mixture of two speakers. Additionally, the proposed two-stage training strategy accelerates convergence, reducing the number of optimization steps required to reach 3 dB SNR by 60%. Overall, our method achieves state-of-the-art performance in the monaural target speaker extraction conditioned on noisy enrollments. Our implementation is available at https://github.com/xu-shitong/TSE-through-Positive-Negative-Enroll .

语音分离噪声处理说话人提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。