针对家庭场景设计专用语音分离模型,提升特定说话人组的识别效果。
TGIF: Talker Group-Informed Familiarization of Target Speaker Extraction
- 用教师模型生成伪干净语音,指导学生模型学习特定说话人群体特征。
- 在家庭对话场景中,相比通用模型,分离性能显著提升。
- 适合家庭设备等小规模、个性化语音处理场景使用。
当前最先进的目标说话人提取(TSE)系统通常设计为泛化于任意混音环境,需要大容量模型作为通用解法。个性化语音增强虽可适配单用户场景,但忽略了仅涉及少数说话人的实际需求,如家庭成员间的语音分离。为此,本文提出说话人群体感知熟悉化(TGIF)的TSE方法,使系统专注于特定说话人群体,但因缺乏纯净语音目标而面临挑战。我们采用知识蒸馏策略,由大型教师模型生成伪纯净语音,指导群体特异性学生模型学习,从而在保持计算效率的同时,有效提取特定说话人群体中的目标说话人。实验表明,该方法在特定说话人群体中优于通用基线模型。TGIF概念凸显了面向多样化真实场景(如家庭设备上的本地化TSE)开发专用解决方案的潜力。
原文摘要 · Abstract (English)
State-of-the-art target speaker extraction (TSE) systems are typically designed to generalize to any given mixing environment, necessitating a model with a large enough capacity as a generalist. Personalized speech enhancement could be a specialized solution that adapts to single-user scenarios, but it overlooks the practical need for customization in cases where only a small number of talkers are involved, e.g., TSE for a specific family. We address this gap with the proposed concept, talker group-informed familiarization (TGIF) of TSE, where the TSE system specializes in a particular group of users, which is challenging due to the inherent absence of a clean speech target. To this end, we employ a knowledge distillation approach, where a group-specific student model learns from the pseudo-clean targets generated by a large teacher model. This tailors the student model to effectively extract the target speaker from the particular talker group while maintaining computational efficiency. Experimental results demonstrate that our approach outperforms the baseline generic models by adapting to the unique speech characteristics of a given speaker group. Our newly proposed TGIF concept underscores the potential of developing specialized solutions for diverse and real-world applications, such as on-device TSE on a family-owned device.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。