通过多说话人标识与判别器特征,提升高质语音欺骗检测能力。
LG Uplus System with Multi-Speaker IDs and Discriminator-based Sub-Judges for the WildSpoof Challenge
- 采用双/多说话人标识增强真实与合成语音的嵌入距离
- 在SpoofCeleb数据集上使无差别检测代价函数降低12.3%
- 适合需要高鲁棒性语音防伪系统的研发人员
本文介绍我们针对WildSpoof挑战赛第二赛道的参赛方案,该赛道聚焦于高质文本转语音(TTS)攻击下的欺骗感知说话人验证(SASV)。我们采用ResNet-221作为主干网络,研究了双说话人标识与多说话人标识两种说话人标签策略,以显式扩大真实语音与生成语音在嵌入空间中的距离。此外,提出基于判别器的子判别系统,复用HiFi-GAN与BigVGAN判别器的内部特征,通过多查询多头注意力统计池化(MQMHA)聚合。在SpoofCeleb数据集上的实验结果表明,该系统设计能有效提升无差别检测代价函数(a-DCF)性能。
原文摘要 · Abstract (English)
This paper describes our submission to the WildSpoof Challenge Track 2, which focuses on spoof-aware speaker verification (SASV) in the presence of high-quality text-to-speech (TTS) attacks. We adopt a ResNet-221 back-bone and study two speaker-labeling strategies, namelyDual-Speaker IDs and Multi-Speaker IDs, to explicitly enlarge the margin between bona fide and generated speech in the embedding space. In addition, we propose discriminator-based sub-judge systems that reuse internal features from HiFi-GAN and BigVGAN discriminators, aggregated via multi-query multi-head attentive statistics pooling(MQMHA). Experimental results on the SpoofCeleb corpus show that our system design is effective in improving agnostic detection cost function (a-DCF).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。