ABC团队为电话语音识别优化了前端嵌入提取器,提升说话人识别性能。
Analysis of ABC Frontend Audio Systems for the NIST-SRE24
- 基于ResNet、ReDimNet和XLS-R设计多种架构,适配不同训练条件。
- 在开放条件下使用11万说话人数据集训练,性能表现优异且鲁棒。
- 提供实用方案,适合追求高精度说话人识别的研究与工程应用。
我们对ABC团队为NIST SRE 2024音频任务开发的嵌入提取器(前端)进行了全面分析。遵循NIST设定的两种场景:仅使用提供的电话录音进行训练(固定条件),或额外加入公开数据(开放条件)。在此约束下,我们构建了针对主流对话式电话语音(CTS)领域的最优说话人嵌入提取器。探索了基于ResNet的不同池化机制、近期提出的ReDimNet架构,以及代表自监督预训练模型家族的XLS-R模型。在开放条件下,我们在包含11万说话人、多语言的VoxBlink2数据集上进行训练。实验表明,VoxBlink训练模型具备良好性能与鲁棒性,为构建顶尖前端提供了可复用的实际方法。
原文摘要 · Abstract (English)
We present a comprehensive analysis of the embedding extractors (frontends) developed by the ABC team for the audio track of NIST SRE 2024. We follow the two scenarios imposed by NIST: using only a provided set of telephone recordings for training (fixed) or adding publicly available data (open condition). Under these constraints, we develop the best possible speaker embedding extractors for the pre-dominant conversational telephone speech (CTS) domain. We explored architectures based on ResNet with different pooling mechanisms, recently introduced ReDimNet architecture, as well as a system based on the XLS-R model, which represents the family of large pre-trained self-supervised models. In open condition, we train on VoxBlink2 dataset, containing 110 thousand speakers across multiple languages. We observed a good performance and robustness of VoxBlink-trained models, and our experiments show practical recipes for developing state-of-the-art frontends for speaker recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。