用语音起点提示,轻松提取目标说话人。
Listen to Extract: Onset-Prompted Target Speaker Extraction
- 将目标说话人录音拼接至混合语音,生成人工语音起始点。
- 在WSJ0-2mix等数据集上达到领先性能。
- 方法极简但效果强,适合实时语音分离场景。
我们提出听觉提取(LExt),一种高效且极其简单的单声道目标说话人分离(TSE)算法。给定目标说话人的注册语音,LExt旨在从该说话人与其他说话人混合的语音中提取目标语音。对于每个混合语音,LExt在波形层面将目标说话人的注册语音拼接至混合信号,并训练深度神经网络(DNN)基于拼接后的信号进行目标语音提取。其原理是:这种方式为目标说话人创建了人工语音起始点,可引导DNN(a)识别目标说话人;(b)学习目标说话人的频谱-时序特征以辅助分离。该简单方法在多个公开的TSE数据集(包括WSJ0-2mix、WHAM!和WHAMR!)上均表现出优异的分离性能。
原文摘要 · Abstract (English)
We propose listen to extract (LExt), a highly-effective while extremely-simple algorithm for monaural target speaker extraction (TSE). Given an enrollment utterance of a target speaker, LExt aims at extracting the target speaker from the speaker's mixed speech with other speakers. For each mixture, LExt concatenates an enrollment utterance of the target speaker to the mixture signal at the waveform level, and trains deep neural networks (DNN) to extract the target speech based on the concatenated mixture signal. The rationale is that, this way, an artificial speech onset is created for the target speaker and it could prompt the DNN (a) which speaker is the target to extract; and (b) spectral-temporal patterns of the target speaker that could help extraction. This simple approach produces strong TSE performance on multiple public TSE datasets including WSJ0-2mix, WHAM! and WHAMR!.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。