arXiv:2409.15742eess.AScs.SD2024-09被引 3

通过优化声纹点与负样本提升家庭环境下的开放集说话人识别准确率。

Enhancing Open-Set Speaker Identification through Rapid Tuning with Speaker Reciprocal Points and Negative Sample

  • 用预训练WavLM+快速微调网络实现高效声纹注册。
  • 引入SRPL+方法,结合真实与合成负样本,提升开放集识别准确率27%。
  • 适合智能家居等多说话人复杂场景的声纹识别系统部署。

本文提出一种新型框架,用于家庭环境中开放集说话人识别(Open-set SID),对实现自然的人机交互至关重要。针对现有模型与分类方法的局限性,该工作结合预训练的WavLM前端与少量样本快速微调的神经网络后端进行声纹注册,并采用任务优化的说话人互惠点学习(SRPL)机制,增强对多个目标说话人的区分能力。进一步提出改进版SRPL+,融合语音合成与真实负样本进行负样本学习,显著提升开放集识别精度。在多种多语言、文本依赖型说话人识别数据集上进行了全面评估,验证了该系统在复杂家庭多说话人场景中的高可用性。实验表明,该方法相比直接使用高效的WavLM base+模型,开放集性能最高提升27%。

原文摘要 · Abstract (English)

This paper introduces a novel framework for open-set speaker identification in household environments, playing a crucial role in facilitating seamless human-computer interactions. Addressing the limitations of current speaker models and classification approaches, our work integrates an pretrained WavLM frontend with a few-shot rapid tuning neural network (NN) backend for enrollment, employing task-optimized Speaker Reciprocal Points Learning (SRPL) to enhance discrimination across multiple target speakers. Furthermore, we propose an enhanced version of SRPL (SRPL+), which incorporates negative sample learning with both speech-synthesized and real negative samples to significantly improve open-set SID accuracy. Our approach is thoroughly evaluated across various multi-language text-dependent speaker recognition datasets, demonstrating its effectiveness in achieving high usability for complex household multi-speaker recognition scenarios. The proposed system enhanced open-set performance by up to 27\% over the directly use of efficient WavLM base+ model.

说话人识别开放集快速微调负样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。