用流匹配方法解决单通道语音分离中的源顺序混乱问题。
Flow Matching-Based Speech Source Separation with Best-of-N Biometric Sampling

- 基于条件流匹配生成有序双源输出,避免源混淆。
- 通过生物特征选优和分块对齐,显著提升识别与验证性能。
- 在多个指标上优于现有方法,适合实际语音应用。
单通道语音分离在真实场景部署中仍面临源顺序模糊、生成模型采样波动以及长音频分段推理困难等问题。本文提出一种基于条件流匹配的方法,生成以混合信号为条件的有序双源输出。训练时使用固定说话人编码器确定源顺序,并在推理阶段重复使用该编码器进行生物特征最优-$N$候选选择和分块级通道对齐。在Libri2Mix基准上,采用SI-SDR、PESQ和ESTOI评估分离质量,并通过cpWER(自动语音识别)和EER(说话人验证)衡量下游影响。结果表明,所提Transformer U-Net变体在客观分离指标上表现优异,且在所有测试设置下均实现最低的语音识别和说话人验证错误率。
原文摘要 · Abstract (English)
Single-channel speech separation remains challenging for real-world deployment due to source permutation ambiguity, sampling variability of generative models, and the difficulty of processing long recordings with chunk-wise inference. We address these issues with a conditional flow-matching-based method that produces an ordered two-source output conditioned on the mixture. A frozen speaker encoder defines the source order during training and is reused at inference for biometric best-of-$N$ candidate selection and chunk-level channel alignment. We evaluate separation quality on Libri2Mix benchmark using SI-SDR, PESQ, and ESTOI, and measure downstream impact using cpWER for automatic speech recognition and EER for speaker verification. The results show that the proposed Transformer U-Net variant is competitive with strong baselines in objective separation metrics and achieves the lowest downstream automatic speech recognition and speaker verification error rates in all evaluated settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。