arXiv:2602.20805cs.SDcs.LG2026-02被引 5

研究说话人身份对语音欺骗检测的影响,提出两种新方法提升检测性能。

Assessing the Impact of Speaker Identity in Speech Spoofing Detection

  • 构建双向任务框架,同时进行说话人识别与欺骗检测,用梯度反转层抑制说话人信息
  • 在四个数据集上平均误报率降低17%,最复杂攻击下最高降48%
  • 适合关注语音安全、对抗攻击防御的研究者和工程师

语音欺骗检测系统通常使用多说话人数据训练,假设生成的特征向量与说话人身份无关,但这一假设尚未验证。本文研究说话人身份对检测系统的影响,提出一种说话人不变的多任务框架(SInMT),包含两种策略:一种在嵌入中建模说话人身份,另一种主动消除该信息。SInMT采用多任务学习,联合进行说话人识别与欺骗检测,并引入梯度反转层实现特征解耦。在四个数据集上的评估显示,所提说话人不变模型相比基线平均等错误率降低17%,对于最复杂的攻击(如A11)最高可降低48%。

原文摘要 · Abstract (English)

Spoofing detection systems are typically trained using diverse recordings from multiple speakers, often assuming that the resulting embeddings are independent of speaker identity. However, this assumption remains unverified. In this paper, we investigate the impact of speaker information on spoofing detection systems. We propose two approaches within our Speaker-Invariant Multi-Task framework, one that models speaker identity within the embeddings and another that removes it. SInMT integrates multi-task learning for joint speaker recognition and spoofing detection, incorporating a gradient reversal layer. Evaluated using four datasets, our speaker-invariant model reduces the average equal error rate by 17% compared to the baseline, with up to 48% reduction for the most challenging attacks (e.g., A11).

语音安全欺骗检测多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。