通过对抗学习消除语言线索,提升语音欺骗检测的跨数据集泛化能力。
Linguistic Bias Mitigation for Spoofing Detection via Gradient Reversal and A Variational Information Bottleneck

- 用教师-学生框架,通过梯度反转抑制语言信息。
- 在九个数据集上相对降低36.2%的等错误率(EER)。
- 适合关注语音安全与模型鲁棒性的研究者。
生成式语音技术的快速发展削弱了语音生物识别的可靠性。当前的欺骗检测器在域内测试中表现良好,但在跨域场景下泛化能力差。我们发现这可能源于语言偏差:模型过度依赖训练数据中的语言线索,导致跨数据集性能下降。为此,提出一种语言不变的欺骗检测框架,采用教师-学生对抗学习机制。语言感知教师模型预先在外部语料上训练,通过梯度反转引导学生检测器最小化语言信息。为防止非语言线索被误删,引入变分信息瓶颈以抑制主要线索。在九个DF Arena数据集上,该方法相较基线实现最高36.2%的相对EER降低。
原文摘要 · Abstract (English)
Rapid advancements in generative speech technology have compromised the reliability of voice biometrics. While current spoofing detectors excel when assessed under in-domain conditions, generalisation to out-of-domain settings is often poor. We show that this can be due to linguistic bias. A reliance on linguistic cues observed in training data can then compromise robustness to cross-data. We propose a linguistic-invariant spoofing detection framework utilizing teacher-student adversarial learning. The linguistic-aware teacher model, pre-trained on linguistic content of an external dataset, guides the student detector via gradient reversal to minimize the linguistic information. To prevent the inadvertent removal of non-linguistic cues, we incorporate a Variational Information Bottleneck to enable suppression of principal cues. Across nine DF Arena datasets, our method achieves up to a 36.2% relative reduction in the EER compare to the baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。