动态嵌入提升目标语音提取的实时性和准确性
DENSE: Dynamic Embedding Causal Target Speech Extraction
- 用自回归机制生成随上下文变化的动态嵌入
- 在短时客观可懂度和信干比上均取得提升
- 适合需要实时处理的复杂语音场景
目标语音提取(TSE)旨在从混合信号中分离出特定目标说话人的语音。现有TSE模型通常使用静态嵌入作为条件来提取目标语音,但静态嵌入难以捕捉语音信号的上下文信息,可能限制模型性能。本文提出一种新型动态嵌入因果目标语音提取模型,通过自回归机制基于已提取语音生成上下文相关的嵌入,实现帧级实时提取。实验表明,该模型在短时客观可懂度(STOI)和信号-失真比(SDR)上均有提升,为复杂场景下的目标语音提取提供了有效解决方案。
原文摘要 · Abstract (English)
Target speech extraction (TSE) focuses on extracting the speech of a specific target speaker from a mixture of signals. Existing TSE models typically utilize static embeddings as conditions for extracting the target speaker's voice. However, the static embeddings often fail to capture the contextual information of the extracted speech signal, which may limit the model's performance. We propose a novel dynamic embedding causal target speech extraction model to address this limitation. Our approach incorporates an autoregressive mechanism to generate context-dependent embeddings based on the extracted speech, enabling real-time, frame-level extraction. Experimental results demonstrate that the proposed model enhances short-time objective intelligibility (STOI) and signal-to-distortion ratio (SDR), offering a promising solution for target speech extraction in challenging scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。