提出首个针对目标说话人识别的语音自监督学习评测基准。
TS-SUPERB: A Target Speech Processing Benchmark for Speech Self-Supervised Learning Models
- 构建四类目标说话人处理任务,用语音嵌入作为定位线索。
- 发现单说话人表现无法预测多说话人环境下的性能。
- 统一编码器联合优化,提升多任务间信息共享效率。
自监督学习(SSL)模型在语音处理任务中取得显著进展,已有多个基准用于验证其效果。然而,以往基准主要聚焦于单说话人场景,对嘈杂、多说话人环境下目标说话人任务的关注较少——这虽更具挑战性,却更贴近实际应用。本文提出目标说话人语音处理通用性能基准TS-SUPERB,涵盖四个广泛认可的目标说话人处理任务,需从语音混合信号中识别目标说话人并提取信息。基准采用注册语音提取的说话人嵌入作为条件线索。实验结果表明,评估SSL模型在目标说话人场景下的表现至关重要,且性能无法简单由相关单说话人任务推断。此外,通过统一的基于SSL的目标语音编码器(包含说话人编码器和提取模块),我们研究了多任务联合优化,利用任务间互信息,验证了其有效性。
原文摘要 · Abstract (English)
Self-supervised learning (SSL) models have significantly advanced speech processing tasks, and several benchmarks have been proposed to validate their effectiveness. However, previous benchmarks have primarily focused on single-speaker scenarios, with less exploration of target-speaker tasks in noisy, multi-talker conditions -- a more challenging yet practical case. In this paper, we introduce the Target-Speaker Speech Processing Universal Performance Benchmark (TS-SUPERB), which includes four widely recognized target-speaker processing tasks that require identifying the target speaker and extracting information from the speech mixture. In our benchmark, the speaker embedding extracted from enrollment speech is used as a clue to condition downstream models. The benchmark result reveals the importance of evaluating SSL models in target speaker scenarios, demonstrating that performance cannot be easily inferred from related single-speaker tasks. Moreover, by using a unified SSL-based target speech encoder, consisting of a speaker encoder and an extractor module, we also investigate joint optimization across TS tasks to leverage mutual information and demonstrate its effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。