用双路径变压器和对抗精修,从混响人声中精准提取目标说话人语音。
Spectron: Target Speaker Extraction using Conditional Transformer with Adversarial Refinement
- 采用条件变压器与双目标训练,增强说话人嵌入一致性与波形编码可逆性。
- 在单一麦克风输入下,性能比基线提升3.12 dB,优于现有方法4.1 dB。
- 适合语音分离、语音增强等场景,尤其适用于无额外数据依赖的任务。
近年来,基于注意力的变压器已成为自然语言处理、计算机视觉、信号处理等多个深度学习领域的标准架构。本文提出一种基于变压器的端到端模型,从单声道多说话人混合音频中提取目标说话人的语音。不同于现有方法,我们引入两个附加目标:强制说话人嵌入一致性与波形编码器可逆性,并联合训练说话人编码器与语音分离器,以更好地捕捉说话人条件嵌入。此外,我们采用多尺度判别器来优化提取语音的感知质量。实验表明,使用双路径变压器作为分离器主干并结合所提出的训练范式,相较于CNN基线模型提升了3.12 dB。最后,与近期先进方法对比显示,本模型在不引入额外数据依赖的情况下,平均性能领先4.1 dB。
原文摘要 · Abstract (English)
Recently, attention-based transformers have become a de facto standard in many deep learning applications including natural language processing, computer vision, signal processing, etc.. In this paper, we propose a transformer-based end-to-end model to extract a target speaker's speech from a monaural multi-speaker mixed audio signal. Unlike existing speaker extraction methods, we introduce two additional objectives to impose speaker embedding consistency and waveform encoder invertibility and jointly train both speaker encoder and speech separator to better capture the speaker conditional embedding. Furthermore, we leverage a multi-scale discriminator to refine the perceptual quality of the extracted speech. Our experiments show that the use of a dual path transformer in the separator backbone along with proposed training paradigm improves the CNN baseline by $3.12$ dB points. Finally, we compare our approach with recent state-of-the-arts and show that our model outperforms existing methods by $4.1$ dB points on an average without creating additional data dependency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。