动态更新语音标识,让目标说话人分离更准。
EvoTSE: Evolving Enrollment for Target Speaker Extraction
- 用历史高置信度估计筛选更新语音标识
- 在跨域场景下性能提升显著,最高增益达12.3%
- 无需额外标注数据,适合实际部署场景
目标说话人提取(TSE)旨在从混合语音中分离出特定说话人的声音,依赖预录的语音标识。尽管TSE避免了盲源分离中的全局排列模糊性,但仍易受说话人混淆影响,即模型错误提取干扰说话人。此外,传统TSE采用静态推理流程,性能受限于固定语音标识的质量。为此,本文提出EvoTSE框架,通过可靠性过滤的检索机制,持续更新语音标识,降低说话人混淆,并放宽对预录语音标识质量的要求,且不依赖额外标注数据。多基准测试表明,EvoTSE在多个场景中实现一致提升,尤其在域外(OOD)测试中表现突出。
原文摘要 · Abstract (English)
Target Speaker Extraction (TSE) aims to isolate a specific speaker's voice from a mixture, guided by a pre-recorded enrollment. While TSE bypasses the global permutation ambiguity of blind source separation, it remains vulnerable to speaker confusion, where models mistakenly extract the interfering speaker. Furthermore, conventional TSE relies on static inference pipeline, where performance is limited by the quality of the fixed enrollment. To overcome these limitations, we propose EvoTSE, an evolving TSE framework in which the enrollment is continuously updated through reliability-filtered retrieval over high-confidence historical estimates. This mechanism reduces speaker confusion and relaxes the quality requirements for pre-recorded enrollment without relying on additional annotated data. Experiments across multiple benchmarks demonstrate that EvoTSE achieves consistent improvements, especially when evaluated on out-of-domain (OOD) scenarios. Our code and checkpoints are available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。