用三类联合建模提升语音伪造检测的可解释性与泛化能力
Integrated Spoofing-Robust Automatic Speaker Verification via a Three-Class Formulation and LLR
- 提出三分类框架统一声纹验证与反伪造,通过类别置信度推导对数似然比
- 在ASVSpoof5上性能相当,在SpoofCeleb上表现更优
- 决策过程可解释性强,无需重训即可适配新评估参数
语音伪造鲁棒声纹验证(SASV)旨在整合声纹验证(ASV)与反伪造(CM)功能。现有方法多采用独立模型得分融合,而部分端到端框架虽将两者集成于单一网络,但多基于双编码器结构,可解释性差,且难以在不重新训练的情况下适应新评估参数。为此,本文提出一种基于三分类形式的统一端到端框架,能够从类别置信度中推导对数似然比(LLR),实现更具可解释性的决策流程。实验表明,在ASVSpoof5数据集上性能与现有方法相当,在SpoofCeleb数据集上取得更优结果。可视化与分析进一步验证了该三分类重构提升了模型可解释性。
原文摘要 · Abstract (English)
Spoofing-robust automatic speaker verification (SASV) aims to integrate automatic speaker verification (ASV) and countermeasure (CM). A popular solution is fusion of independent ASV and CM scores. To better modeling SASV, some frameworks integrate ASV and CM within a single network. However, these solutions are typically bi-encoder based, offer limited interpretability, and cannot be readily adapted to new evaluation parameters without retraining. Based on this, we propose a unified end-to-end framework via a three-class formulation that enables log-likelihood ratio (LLR) inference from class logits for a more interpretable decision pipeline. Experiments show comparable performance to existing methods on ASVSpoof5 and better results on SpoofCeleb. The visualization and analysis also prove that the three-class reformulation provides more interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。