arXiv:2409.15799eess.AScs.SD2024-09被引 30

WeSep工具包让目标说话人分离更易用,支持灵活建模和快速部署。

WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction

  • 支持灵活的目标说话人建模,可适应不同场景需求
  • 内置实时数据生成机制,提升训练效率与泛化能力
  • 适合语音识别、助听器等实际应用开发人员使用

目标说话人提取(TSE)旨在从多人重叠的语音中分离出特定目标说话人的声音,是典型鸡尾酒会问题的解决方案。近年来,由于其在个性化人机交互、助听设备以及语音识别、说话人识别等后续任务中的重要作用,受到越来越多关注。然而,目前缺乏可用于开箱即用的开源工具包或预训练模型。本文提出WeSep,一个面向研究与实际应用的TSE工具包,具备灵活的目标说话人建模、可扩展的数据管理、高效的在线数据模拟、结构化配置模板及部署支持。该工具包已公开发布于 url{https://github.com/wenet-e2e/WeSep}。

原文摘要 · Abstract (English)

Target speaker extraction (TSE) focuses on isolating the speech of a specific target speaker from overlapped multi-talker speech, which is a typical setup in the cocktail party problem. In recent years, TSE draws increasing attention due to its potential for various applications such as user-customized interfaces and hearing aids, or as a crutial front-end processing technologies for subsequential tasks such as speech recognition and speaker recongtion. However, there are currently few open-source toolkits or available pre-trained models for off-the-shelf usage. In this work, we introduce WeSep, a toolkit designed for research and practical applications in TSE. WeSep is featured with flexible target speaker modeling, scalable data management, effective on-the-fly data simulation, structured recipes and deployment support. The toolkit is publicly avaliable at \url{https://github.com/wenet-e2e/WeSep.}

说话人分离语音处理工具包

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。