arXiv:2511.01652eess.AScs.SD2025-11中稿 · APSIPA ASC 2025被引 1

用预训练模型提升多语言语音分离效果

Leveraging Language Information for Target Language Extraction

  • 利用语音预训练模型提供语言先验知识
  • 英文和德文分离性能分别提升1.22和1.12 dB
  • 首个公开的多语言语音目标提取数据集

目标语言提取旨在从包含多名说不同语言者的混合语音波形中提取特定语言的语音。人类听觉系统凭借对特定语言的认知可高效完成此任务,但传统系统因缺乏此类先验知识而表现受限。语音预训练模型从大规模真实语料中捕捉丰富的语言与语音表征,可弥补这一不足。本文提出一种新型端到端框架,利用预训练模型中的语言知识引导提取模型更准确地捕捉目标语言特征,从而提升分离质量。为验证方法有效性,我们构建了首个公开的多语言目标语言提取数据集。实验表明,该方法在含英、德双语的混合语音中,分别实现1.22 dB和1.12 dB的SI-SNR提升。

原文摘要 · Abstract (English)

Target Language Extraction aims to extract speech in a specific language from a mixture waveform that contains multiple speakers speaking different languages. The human auditory system is adept at performing this task with the knowledge of the particular language. However, the performance of the conventional extraction systems is limited by the lack of this prior knowledge. Speech pre-trained models, which capture rich linguistic and phonetic representations from large-scale in-the-wild corpora, can provide this missing language knowledge to these systems. In this work, we propose a novel end-to-end framework to leverage language knowledge from speech pre-trained models. This knowledge is used to guide the extraction model to better capture the target language characteristics, thereby improving extraction quality. To demonstrate the effectiveness of our proposed approach, we construct the first publicly available multilingual dataset for Target Language Extraction. Experimental results show that our method achieves improvements of 1.22 dB and 1.12 dB in SI-SNR for English and German extraction, respectively, from mixtures containing both languages.

语音分离多语言预训练模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。