用流匹配方法实现精准目标说话人分离,高效且无需复杂预训练组件。
FlowTSE: Target Speaker Extraction with Flow Matching
- 基于条件流匹配,直接从语音混合信号中提取目标说话人声音。
- 在标准数据集上表现媲美甚至超越现有强基线模型。
- 适合需要高精度语音分离的场景,如会议记录与语音增强。
目标说话人提取(TSE)旨在利用说话人注册信息,从语音混合信号中分离出特定说话人的纯净语音。尽管现有方法多为判别式,近年生成式方法已取得显著进展,但针对TSE的生成式方法仍处于探索阶段,多数方案依赖复杂流水线和预训练组件,带来较高计算开销。本文提出FlowTSE,一种基于条件流匹配的简洁有效TSE方法。模型接收注册语音样本与混合语音信号,均以梅尔频谱图表示,目标是重建目标说话人的干净语音。此外,针对相位重建至关重要的任务,我们设计了一种新型声码器,其条件输入为混合信号的复数STFT,从而提升相位估计性能。在标准TSE基准上的实验结果表明,FlowTSE在性能上达到或超过多个强基线模型。
原文摘要 · Abstract (English)
Target speaker extraction (TSE) aims to isolate a specific speaker's speech from a mixture using speaker enrollment as a reference. While most existing approaches are discriminative, recent generative methods for TSE achieve strong results. However, generative methods for TSE remain underexplored, with most existing approaches relying on complex pipelines and pretrained components, leading to computational overhead. In this work, we present FlowTSE, a simple yet effective TSE approach based on conditional flow matching. Our model receives an enrollment audio sample and a mixed speech signal, both represented as mel-spectrograms, with the objective of extracting the target speaker's clean speech. Furthermore, for tasks where phase reconstruction is crucial, we propose a novel vocoder conditioned on the complex STFT of the mixed signal, enabling improved phase estimation. Experimental results on standard TSE benchmarks show that FlowTSE matches or outperforms strong baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。