通过小波提示调优与多模型融合,提升语音伪造检测与身份验证的联合防御能力。
Spoofing-Aware Speaker Verification via Wavelet Prompt Tuning and Multi-Model Ensembles
- 采用小波提示调优的XLSR-AASIST检测伪造,再用多模型集成验证身份
- 在域内数据上伪造检测误报率低至0.16%,跨域测试仍面临挑战
- 适合关注语音安全、对抗生成攻击的研究者与开发者
本文介绍乌兹堡大学-苏黎世联邦理工学院(UZH-CL)提交至WildSpoof 2026挑战赛的SASV赛道系统。该挑战要求同时验证说话人身份与音频真实性,以应对生成式伪造攻击。我们提出一种级联式防伪造说话人验证框架,融合小波提示调优的XLSR-AASIST反伪造模块与多模型集成方案。ASV组件采用ResNet34、ResNet293和WavLM-ECAPA-TDNN架构,经Z-score归一化后进行得分平均。系统在VoxCeleb2和SpoofCeleb数据集上训练,获得宏观a-DCF为0.2017,SASV-EER为2.08%。在域内数据上,伪造检测EER达0.16%;但在未见数据集ASVspoof5上的表现揭示了跨域泛化能力的显著挑战。
原文摘要 · Abstract (English)
This paper describes the UZH-CL system submitted to the SASV section of the WildSpoof 2026 challenge. The challenge focuses on the integrated defense against generative spoofing attacks by requiring the simultaneous verification of speaker identity and audio authenticity. We proposed a cascaded Spoofing-Aware Speaker Verification framework that integrates a Wavelet Prompt-Tuned XLSR-AASIST countermeasure with a multi-model ensemble. The ASV component utilizes the ResNet34, ResNet293, and WavLM-ECAPA-TDNN architectures, with Z-score normalization followed by score averaging. Trained on VoxCeleb2 and SpoofCeleb, the system obtained a Macro a-DCF of 0.2017 and a SASV EER of 2.08%. While the system achieved a 0.16% EER in spoof detection on the in-domain data, results on unseen datasets, such as the ASVspoof5, highlight the critical challenge of cross-domain generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。