提出可学习的声纹验证归一化方法,提升不同语音的校准精度。
Trainable Adaptive Score Normalization for Automatic Speaker Verification
- 用可训练的伪说话人嵌入构建评分参照组
- 在VoxCeleb1-O上等错误率降低4.11%,最小检测代价函数降10.62%
- 适用于多语言场景,鲁棒性强
自适应S-归一化(AS-norm)通过使用与目标说话人相似的伪造者得分来校准声纹验证(ASV)得分,但其无学习过程,难以针对不同测试语音提供合适的正则化强度。为此,本文提出可训练的AS-norm(TAS-norm),引入可学习的伪造者嵌入(LIEs)构成参照组。LIEs初始设定为训练数据集中伪造者说话人的表征,并通过模拟声纹验证过程进行微调。微调时采用边界惩罚机制,在选择高分伪造者嵌入时防止非伪造者被选中。实验基于ECAPA-TDNN模型,在VoxCeleb1-O测试集上,TAS-norm相比标准AS-norm分别实现4.11%和10.62%的相对等错误率与最小检测代价函数改进。进一步在波斯语和汉语数据集上验证了其跨语言有效性,证明该方法具有良好的泛化能力。
原文摘要 · Abstract (English)
Adaptive S-norm (AS-norm) calibrates automatic speaker verification (ASV) scores by normalizing them utilize the scores of impostors which are similar to the input speaker. However, AS-norm does not involve any learning process, limiting its ability to provide appropriate regularization strength for various evaluation utterances. To address this limitation, we propose a trainable AS-norm (TAS-norm) that leverages learnable impostor embeddings (LIEs), which are used to compose the cohort. These LIEs are initialized to represent each speaker in a training dataset consisting of impostor speakers. Subsequently, LIEs are fine-tuned by simulating an ASV evaluation. We utilize a margin penalty during top-scoring IEs selection in fine-tuning to prevent non-impostor speakers from being selected. In our experiments with ECAPA-TDNN, the proposed TAS-norm observed 4.11% and 10.62% relative improvement in equal error rate and minimum detection cost function, respectively, on VoxCeleb1-O trial compared with standard AS-norm without using proposed LIEs. We further validated the effectiveness of the TAS-norm on additional ASV datasets comprising Persian and Chinese, demonstrating its robustness across different languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。