通过参考音频训练提升语音反欺骗检测效果,推理时可忽略参考音频
RAT: Reference-Augmented Training for ASV Anti-Spoofing
- 利用参考录音进行训练,使模型对语音伪造更具鲁棒性
- 在ASVspoof 5上达到2.57% EER和0.074 minDCF的顶尖性能
- 适合需要高精度语音反欺骗的系统开发者使用
我们提出一种基于说话人参考录音的欺骗防御架构,但发现模型在推理时会忽略参考信息。令人意外的是,训练中引入参考通道能促使模型产生对欺骗的不变性,即使推理时参考缺失或不匹配也能提升深度伪造检测能力。基于此,我们提出参考增强训练(RAT)策略。RAT在单检测器设置下超越传统单语句基线,即便推理时将参考替换为零向量也表现优异。严谨分析表明,优化过程迅速弱化参考贡献,使推理几乎不受参考通道影响。采用RAT,我们在ASVspoof 5基准上实现2.57% EER和0.074 minDCF的当前最优结果,优于大型集成系统。
原文摘要 · Abstract (English)
We introduce a spoofing countermeasure architecture conditioned on speaker-reference recordings, but observe that it converges to a solution that effectively ignores the reference during inference. Surprisingly, training with a reference channel induces invariance that improves deepfake detection, even when the reference is absent or mismatched during inference. Based on this observation, we propose a Reference-Augmented Training (RAT) strategy. RAT yields improved detection performance compared to single-utterance baselines, even when the reference recording is replaced with a zero vector at inference. Through rigorous analysis, we demonstrate that the optimization process rapidly diminishes the reference contributions, leading to inference largely independent of the reference channel. Using RAT, we achieve state-of-the-art 2.57% EER and 0.074 minDCF on the ASVspoof 5 benchmark with a single detector, surpassing even large ensemble systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。