一歩で高品質な目標話者抽出を実現する新アーキテクチャ
MeanFlow-TSE: One-Step Generative Target Speaker Extraction with Mean Flow
- 流体理論に基づく平均フロー学習で、1ステップ生成を可能に
- リブリ2ミックスデータで既存手法を上回る音声分離性能
- リアルタイム処理向けに低遅延で高精度な抽出が可能
目标说话人提取(TSE)旨在利用参考语音等辅助信息从多人混合语音中分离出指定说话人的声音。尽管扩散模型和流匹配模型的进展提升了TSE性能,但这些方法通常需要多步采样,限制了其在低延迟场景中的应用。本文提出MeanFlow-TSE,一种基于平均流目标训练的一步生成式TSE框架,可在无需迭代优化的情况下实现快速且高质量的生成。基于AD-FlowTSE范式,本方法定义了由混音比例(MR)控制的背景与目标声源之间的流变换。在Libri2Mix语料库上的实验表明,该方法在分离质量和感知指标方面均优于现有的基于扩散和流匹配的TSE模型,且仅需单次推理步骤。结果表明,基于平均流引导的一步生成为实时目标说话人提取提供了一种高效有效的替代方案。代码已开源:https://github.com/rikishimizu/MeanFlow-TSE。
原文摘要 · Abstract (English)
Target speaker extraction (TSE) aims to isolate a desired speaker's voice from a multi-speaker mixture using auxiliary information such as a reference utterance. Although recent advances in diffusion and flow-matching models have improved TSE performance, these methods typically require multi-step sampling, which limits their practicality in low-latency settings. In this work, we propose MeanFlow-TSE, a one-step generative TSE framework trained with mean-flow objectives, enabling fast and high-quality generation without iterative refinement. Building on the AD-FlowTSE paradigm, our method defines a flow between the background and target source that is governed by the mixing ratio (MR). Experiments on the Libri2Mix corpus show that our approach outperforms existing diffusion- and flow-matching-based TSE models in separation quality and perceptual metrics while requiring only a single inference step. These results demonstrate that mean-flow-guided one-step generation offers an effective and efficient alternative for real-time target speaker extraction. Code is available at https://github.com/rikishimizu/MeanFlow-TSE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。