arXiv:2606.08678cs.SDcs.LG2026-06

让语音伪造检测模型摆脱说话人特征干扰,提升跨场景泛化能力。

Speaker-Invariant Representation Learning for Spoofing Detection via Gradient Reversal and A Variational Information Bottleneck

论文配图:Speaker-Invariant Representation Learning for Spoofing Detection via Gradient Reversal and A Variational Information Bottleneck
图 1 · 摘自论文原文
  • 用教师-学生框架+梯度反转层剥离说话人身份信息。
  • 在9个数据集上将误报率相对降低25.7%。
  • 无需说话人标签,适合真实环境中复杂语音场景。

先进的生成式语音技术正威胁语音生物识别的可靠性。尽管现有伪造检测系统在同域条件下表现优异,但在跨域场景中泛化能力普遍较差。本文发现,问题根源在于说话人偏差——模型过度依赖个体语音特征而非伪造痕迹。为此,我们提出一种无需说话人标签的师生框架,利用预训练的说话人识别教师模型通过梯度反转层指导学生模型,同时引入变分信息瓶颈(Variational Information Bottleneck)以平衡身份线索抑制与伪造线索保留。在九个数据集上的评估表明,该方法相比MHFA基线将等错误率(EER)降低了25.7%。

原文摘要 · Abstract (English)

Sophisticated generative speech technology can undermined the reliability of voice biometrics. While spoofing detection systems excel when assessed under in-domain conditions, generalisation to out-of-domain settings is often poor. In this paper, we show that such issues could be caused by speaker bias, where models learn individual voice traits rather than markers of manipulation or generation. We propose a teacher-student framework for speaker-invariant spoofing detection that disentangles identity without requiring speaker labels. We leverage a pre-trained speaker recognition teacher to guide a student model via a gradient reversal layer. To control the balance between suppressing cues related to voice identity with the preservation of those related to spoofing detection, we integrate a Variational Information Bottleneck. Evaluations across nine datasets show our model achieves a 25.7% relative reduction to the EER compared to the MHFA baseline.

语音伪造身份无关信息瓶颈对抗学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。