提出联合优化说话人与欺骗检测的模块化框架,提升语音认证鲁棒性。
Joint Optimization of Speaker and Spoof Detectors for Spoofing-Robust Automatic Speaker Verification
- 用可训练后端融合说话人与欺骗检测输出,实现模块化联合优化
- 非线性分数融合使a-DCF降低至0.196,SPF-EER降至7.6%
- 适合需要高鲁棒性和可解释性的实际语音认证系统
抗欺骗说话人验证(SASV)在对抗环境下同时完成说话人识别与欺骗检测。现有系统多采用独立训练子系统,在嵌入、得分或决策层进行融合。本文保留两个子系统的模块化结构,通过可训练后端分类器整合其输出,并探索直接以最新提出的SASV评估指标a-DCF为优化目标的训练方法。在ASVspoof 5数据集上的实验表明:(i) 非线性分数融合持续优于线性融合;(ii) 结合加权余弦评分(weighted cosine scoring)用于说话人检测、SSL-AASIST用于欺骗检测,达到当前最优性能,最小a-DCF为0.196,SPF-EER为7.6%。这些结果凸显了模块化设计、校准融合与任务对齐优化对构建鲁棒且可解释SASV系统的重要性。
原文摘要 · Abstract (English)
Spoofing-robust speaker verification (SASV) combines the tasks of speaker and spoof detection to authenticate speakers under adversarial settings. Many SASV systems rely on fusion of speaker and spoof cues at embedding, score or decision levels, based on independently trained subsystems. In this study, we respect similar modularity of the two subsystems, by integrating their outputs using trainable back-end classifiers. In particular, we explore various approaches for directly optimizing the back-end for the recently-proposed SASV performance metric (a-DCF) as a training objective. Our experiments on the ASVspoof 5 dataset demonstrate two important findings: (i) nonlinear score fusion consistently improves a-DCF over linear fusion, and (ii) the combination of weighted cosine scoring for speaker detection with SSL-AASIST for spoof detection achieves state-of-the-art performance, reducing min a-DCF to 0.196 and SPF-EER to 7.6%. These contributions highlight the importance of modular design, calibrated integration, and task-aligned optimization for advancing robust and interpretable SASV systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。