提出可处理不同重叠率的语音提取模型,提升真实场景下的分离效果。
VorTEX: Various overlap ratio for Target speech EXtraction
- 设计解耦自适应多分支融合结构,分离主提取与辅助正则路径。
- 在20%-100%重叠率下实现最高分离保真度(20%时5.50 dB,100%时2.04 dB)。
- 构建新诊断指标SuRE,发现并避免传统方法的抑制伪影问题。
目标语音提取(TSE)旨在从混合语音中恢复特定说话人的声音。尽管近期基于文本提示的方法展现出潜力,但大多数研究假设混合信号完全重叠,限制了对实际重叠比率下行为的理解。本文提出VorTEX(Various overlap ratio for Target speech EXtraction),一种基于文本提示的TSE架构,采用解耦自适应多分支(DAM)融合模块,将主提取路径与辅助正则化路径分离。为支持可控分析,我们构建了PORTE数据集,包含覆盖0%至100%重叠比率的双说话人混合语音。同时提出能量抑制比(SuRE)作为诊断指标,检测传统评估方法未捕捉的抑制行为。实验表明,现有模型在重叠情况下存在抑制或残留干扰,而VorTEX在20%-100%重叠范围内达到最高分离保真度(如20%时为5.50 dB,100%时为2.04 dB),且保持零SuRE,表明其在无抑制伪影前提下稳健提取语音。
原文摘要 · Abstract (English)
Target speech extraction (TSE) aims to recover a target speaker's voice from a mixture. While recent text-prompted approaches have shown promise, most approaches assume fully overlapped mixtures, limiting insight into behavior across realistic overlap ratios. We introduce VorTEX (Various overlap ratio for Target speech EXtraction), a text-prompted TSE architecture with a Decoupled Adaptive Multi-branch (DAM) Fusion block that separates primary extraction from auxiliary regularization pathways. To enable controlled analysis, we construct PORTE, a two-speaker dataset spanning overlap ratios from 0% to 100%. We further propose Suppression Ratio on Energy (SuRE), a diagnostic metric that detects suppression behavior not captured by conventional measures. Experiments show that existing models exhibit suppression or residual interference under overlap, whereas VorTEX achieves the highest separation fidelity across 20-100% overlap (e.g., 5.50 dB at 20% and 2.04 dB at 100%) while maintaining zero SuRE, indicating robust extraction without suppression-driven artifacts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。