用关键词定位说话人,无需提前录音就能提取目标语音。
Detect, Attend and Extract: Keyword Guided Target Speaker Extraction
- 通过关键词检测、注意力聚焦和语音提取三步定位目标说话人。
- 在无预录语音场景下性能优于传统依赖录入的系统。
- 适合实际应用中无法获取清晰录音的场景,如会议记录、监控等。
目标说话人分离(TSE)旨在从多说话人混合语音中提取特定目标说话人的语音。传统TSE系统主要依赖预录入的语音特征来识别目标说话人,但在许多实际场景中,干净的录入语音难以获取,限制了现有方法的应用。本文提出DAE-TSE框架,通过目标说话人说出的特定关键词(即部分转录文本)来指定目标说话人。该方法遵循‘检测-注意-提取’(DAE)范式:首先检测给定关键词是否存在,然后根据关键词内容关注对应说话人,最后提取目标语音。实验表明,DAE-TSE在无清洁录入语音条件下表现优于标准TSE系统。据我们所知,这是首个将部分转录文本作为提示来指定目标说话人的研究,为真实场景提供了灵活且实用的解决方案。代码与演示页面已公开。
原文摘要 · Abstract (English)
Target speaker extraction (TSE) aims to extract the speech of a target speaker from mixtures containing multiple competing speakers. Conventional TSE systems predominantly rely on speaker cues, such as pre-enrolled speech, to identify and isolate the target speaker. However, in many practical scenarios, clean enrollment utterances are unavailable, limiting the applicability of existing approaches. In this work, we propose DAE-TSE, a keyword-guided TSE framework that specifies the target speaker through distinct keywords they utter. By leveraging keywords (i.e., partial transcriptions) as cues, our approach provides a flexible and practical alternative to enrollment-based TSE. DAE-TSE follows the Detect-Attend-Extract (DAE) paradigm: it first detects the presence of the given keywords, then attends to the corresponding speaker based on the keyword content, and finally extracts the target speech. Experimental results demonstrate that DAE-TSE outperforms standard TSE systems that rely on clean enrollment speech. To the best of our knowledge, this is the first study to utilize partial transcription as a cue for specifying the target speaker in TSE, offering a flexible and practical solution for real-world scenarios. Our code and demo page are now publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。