arXiv:2409.07841cs.SDcs.LG2024-09被引 15

用离散令牌和语言模型实现高质量目标说话人提取

TSELM: Target Speaker Extraction using Discrete Tokens and Language Models

  • 用WavLM的多层离散特征作为输入,结合交叉注意力融合目标说话人信息
  • 将音频生成转为分类任务,通过交叉熵损失建模输出令牌分布
  • 音质表现优秀,适合需要高保真语音分离的场景

我们提出TSELM,一种利用离散令牌和语言模型进行目标说话人提取的新方法。TSELM采用WavLM的多层离散化表示作为输入令牌,并通过交叉注意力机制整合目标说话人信息。语言模型用于捕捉序列依赖关系,同时使用可扩展的HiFi-GAN从令牌重建音频。通过交叉熵损失,TSELM建模输出令牌的概率分布,从而将复杂的音频生成回归问题转化为分类任务。实验结果表明,TSELM在语音质量上表现优异,在语音可懂度方面达到相当水平。

原文摘要 · Abstract (English)

We propose TSELM, a novel target speaker extraction network that leverages discrete tokens and language models. TSELM utilizes multiple discretized layers from WavLM as input tokens and incorporates cross-attention mechanisms to integrate target speaker information. Language models are employed to capture the sequence dependencies, while a scalable HiFi-GAN is used to reconstruct the audio from the tokens. By applying a cross-entropy loss, TSELM models the probability distribution of output tokens, thus converting the complex regression problem of audio generation into a classification task. Experimental results show that TSELM achieves excellent results in speech quality and comparable results in speech intelligibility.

说话人分离离散令牌语音生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。