用真实录音训练模型,从混音中分离出说话人自己的声音。
Cross-Talk Speech Reduction, by Separation, for Separation

- 通过真实近场与远场录音对,直接训练模型去除近场混音中的串音
- 在CHiME-6数据集上语音识别效果超越所有此前挑战赛提交方案
- 无需模拟数据,适合真实场景下的语音分离任务
在对话语音分离与识别任务中,通常在训练数据收集时为每位说话人配备近场麦克风以捕获近场混合信号,同时使用远场麦克风记录远场混合信号。每个近场混合信号对佩戴者而言能量较高,可直观作为弱监督信号,用于直接在真实远场信号上训练远场语音分离模型。然而,这些信号因包含其他说话人的强烈串音及背景噪声,清洁度不足。为此,我们提出交叉通话抑制(CTR)任务,以及一种新方法CTRnet,可在真实录制的近场与远场混合信号对上直接训练,实现从近场混合信号中分离出佩戴者语音。基于CTRnet,进一步提出伪标签远场语音分离(PuLSS),利用CTRnet估计的干净语音作为伪标签,训练远场混合信号分离模型。该框架关键优势在于,CTRnet和PuLSS均可在目标域的真实录音上训练,解决了仅用模拟数据训练导致的泛化差距问题。在CHiME-6数据集上,该框架在理想与估计说话人分段条件下均达到当前最优的语音识别性能,超越所有CHiME-{7,8}挑战赛提交结果。据我们所知,这是首个在真实“野外”对话语音数据上显著优于引导源分离的神经语音分离方法。
原文摘要 · Abstract (English)
In conversational speech separation and recognition tasks, close-talk microphones are typically attached to each speaker during training data collection to capture near-field, close-talk mixture signals, in addition to using far-field microphones to record far-field mixture signals. Each such close-talk mixture exhibits a reasonably high energy level for the wearer and could intuitively serve as weak supervision for training far-field speech separation models directly on real-recorded far-field signals. However, they are not sufficiently clean for this purpose, as they often contain strong cross-talk speech from other speakers in addition to background noise. To address this, we propose cross-talk reduction (CTR), a task aiming to isolate the wearer's speech from each close-talk mixture, and a novel method called CTRnet, which can be trained directly on real-recorded pairs of close-talk and far-field mixtures to accomplish CTR. Building on CTRnet, we further propose pseudo-label based far-field speech separation (PuLSS), which uses CTRnet's estimated clean speech as pseudo-labels to train models for separating far-field mixtures. A key advantage of the proposed framework is that both CTRnet and PuLSS can be trained on real-recorded data from the target domain, addressing the generalization gap commonly observed when models are trained exclusively on simulated data. On the CHiME-6 dataset, our framework achieves state-of-the-art ASR performance under both oracle and estimated speaker diarization, surpassing all CHiME-{7,8} challenge submissions. To our knowledge, it is the first neural speech separation method that substantially outperforms guided source separation on real conversational "speech-in-the-wild" data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。