arXiv:2505.21805cs.SDeess.AS2025-05被引 2

通过生成伪说话人提升语音分离模型的辨识能力

An Investigation on Speaker Augmentation for End-to-End Speaker Extraction

  • 在时域上重采样缩放,生成多样伪说话人特征
  • 在WSJ0-2Mix和LibriMix上降低目标混淆率
  • 可与度量学习结合,适合语音分离研究者

目标混淆(即偶尔切换到非目标说话人)是端到端说话人提取(E2E-SE)系统的关键挑战。我们认为该问题主要源于说话人嵌入的泛化性和区分性不足,提出一种简单而有效的说话人增强策略。具体而言,设计了一种时域重采样与缩放流程,在保留其他语音特性的同时改变说话人特征,生成多种伪说话人,以建立更具泛化性的说话人嵌入空间;同时,特定于说话人特征的增强产生困难样本,迫使模型关注真实说话人特征。在WSJ0-2Mix和LibriMix上的实验表明,该方法有效缓解了目标混淆并提升了提取性能。此外,该方法可与度量学习结合,实现进一步增益。

原文摘要 · Abstract (English)

Target confusion, defined as occasional switching to non-target speakers, poses a key challenge for end-to-end speaker extraction (E2E-SE) systems. We argue that this problem is largely caused by the lack of generalizability and discrimination of the speaker embeddings, and introduce a simple yet effective speaker augmentation strategy to tackle the problem. Specifically, we propose a time-domain resampling and rescaling pipeline that alters speaker traits while preserving other speech properties. This generates a variety of pseudo-speakers to help establish a generalizable speaker embedding space, while the speaker-trait-specific augmentation creates hard samples that force the model to focus on genuine speaker characteristics. Experiments on WSJ0-2Mix and LibriMix show that our method mitigates the target confusion and improves extraction performance. Moreover, it can be combined with metric learning, another effective approach to address target confusion, leading to further gains.

语音分离说话人增强嵌入学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。