一拍即合:用单步生成实现高保真说话人提取
AlphaFlowTSE: One-Step Generative Target Speaker Extraction via Conditional AlphaFlow
- 基于无雅可比向量积的AlphaFlow目标,单步完成说话人分离
- 在Libri2Mix和REAL-T上提升语音相似度与实际场景泛化能力
- 适合追求低延迟、高稳定性的语音识别前端应用
在目标说话人提取(TSE)中,需利用一段短参考语音从多人混叠语音中恢复目标语音。近期基于扩散模型与流匹配生成器的研究提升了目标语音保真度,但多步采样导致延迟较高,而单步方法常依赖混合比例相关的时序坐标,在真实对话中不可靠。本文提出AlphaFlowTSE,一种基于无雅可比向量积的条件生成模型,通过学习从混合信号出发的目标轨迹上的平均速度迁移,无需额外的混合比例预测,并结合区间一致性师生目标稳定训练过程。在Libri2Mix与REAL-T数据集上的实验表明,AlphaFlowTSE显著提升目标说话人相似度及真实混叠场景下的泛化性能,有利于下游自动语音识别(ASR)任务。
原文摘要 · Abstract (English)
In target speaker extraction (TSE), we aim to recover target speech from a multi-talker mixture using a short enrollment utterance as reference. Recent studies on diffusion and flow-matching generators have improved target-speech fidelity. However, multi-step sampling increases latency, and one-step solutions often rely on a mixture-dependent time coordinate that can be unreliable for real-world conversations. We present AlphaFlowTSE, a one-step conditional generative model trained with a Jacobian-vector product (JVP)-free AlphaFlow objective. AlphaFlowTSE learns mean-velocity transport along a mixture-to-target trajectory starting from the observed mixture, eliminating auxiliary mixing-ratio prediction, and stabilizes training by combining flow matching with an interval-consistency teacher-student target. Experiments on Libri2Mix and REAL-T confirm that AlphaFlowTSE improves target-speaker similarity and real-mixture generalization for downstream automatic speech recognition (ASR).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。