通过自适应步长实现高效精准的语音目标提取
Adaptive Deterministic Flow Matching for Target Speaker Extraction
- 基于流匹配,按混响比动态调整反向生成路径
- 仅需一步即可达成强效果,且混响比估计提升语音精度
- 适合追求高效高保真语音分离的工业应用
生成式目标说话人提取(TSE)方法通常比预测模型生成更自然的输出。现有基于扩散或流匹配(FM)的方法通常采用固定数量的反向步骤和固定步长。本文提出自适应判别流匹配目标说话人提取(AD-FlowTSE),通过自适应步长提取目标语音。在流匹配框架下,不同于以往从混合信号(或标准先验)到干净语音分布的传输,我们定义背景与源语音之间的流,由混合比例(MR)控制。该设计支持MR感知初始化,使模型从背景-源语音轨迹上的自适应点开始,而非对所有噪声水平使用统一反向调度。实验表明,AD-FlowTSE仅用一步即可实现优异的TSE性能,且结合辅助MR估计可进一步提升目标语音准确性。结果表明,将传输路径与混合成分对齐,并根据噪声条件自适应调整步长,能实现高效且准确的语音提取。
原文摘要 · Abstract (English)
Generative target speaker extraction (TSE) methods often produce more natural outputs than predictive models. Recent work based on diffusion or flow matching (FM) typically relies on a small, fixed number of reverse steps with a fixed step size. We introduce Adaptive Discriminative Flow Matching TSE (AD-FlowTSE), which extracts the target speech using an adaptive step size. We formulate TSE within the FM paradigm but, unlike prior FM-based speech enhancement and TSE approaches that transport between the mixture (or a normal prior) and the clean-speech distribution, we define the flow between the background and the source, governed by the mixing ratio (MR) of the source and background that creates the mixture. This design enables MR-aware initialization, where the model starts at an adaptive point along the background-source trajectory rather than applying the same reverse schedule across all noise levels. Experiments show that AD-FlowTSE achieves strong TSE with as few as a single step, and that incorporating auxiliary MR estimation further improves target speech accuracy. Together, these results highlight that aligning the transport path with the mixture composition and adapting the step size to noise conditions yields efficient and accurate TSE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。