arXiv:2504.04512eess.AS2025-04中稿 · ICASSP'25被引 3

提出可学习的声纹验证归一化方法,提升不同语音的校准精度。

Trainable Adaptive Score Normalization for Automatic Speaker Verification

  • 用可训练的伪说话人嵌入构建评分参照组
  • 在VoxCeleb1-O上等错误率降低4.11%,最小检测代价函数降10.62%
  • 适用于多语言场景,鲁棒性强

自适应S-归一化(AS-norm)通过使用与目标说话人相似的伪造者得分来校准声纹验证(ASV)得分,但其无学习过程,难以针对不同测试语音提供合适的正则化强度。为此,本文提出可训练的AS-norm(TAS-norm),引入可学习的伪造者嵌入(LIEs)构成参照组。LIEs初始设定为训练数据集中伪造者说话人的表征,并通过模拟声纹验证过程进行微调。微调时采用边界惩罚机制,在选择高分伪造者嵌入时防止非伪造者被选中。实验基于ECAPA-TDNN模型,在VoxCeleb1-O测试集上,TAS-norm相比标准AS-norm分别实现4.11%和10.62%的相对等错误率与最小检测代价函数改进。进一步在波斯语和汉语数据集上验证了其跨语言有效性,证明该方法具有良好的泛化能力。

原文摘要 · Abstract (English)

Adaptive S-norm (AS-norm) calibrates automatic speaker verification (ASV) scores by normalizing them utilize the scores of impostors which are similar to the input speaker. However, AS-norm does not involve any learning process, limiting its ability to provide appropriate regularization strength for various evaluation utterances. To address this limitation, we propose a trainable AS-norm (TAS-norm) that leverages learnable impostor embeddings (LIEs), which are used to compose the cohort. These LIEs are initialized to represent each speaker in a training dataset consisting of impostor speakers. Subsequently, LIEs are fine-tuned by simulating an ASV evaluation. We utilize a margin penalty during top-scoring IEs selection in fine-tuning to prevent non-impostor speakers from being selected. In our experiments with ECAPA-TDNN, the proposed TAS-norm observed 4.11% and 10.62% relative improvement in equal error rate and minimum detection cost function, respectively, on VoxCeleb1-O trial compared with standard AS-norm without using proposed LIEs. We further validated the effectiveness of the TAS-norm on additional ASV datasets comprising Persian and Chinese, demonstrating its robustness across different languages.

声纹验证可学习归一化多语言嵌入学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。