arXiv:2603.18657cs.LG2026-03被引 1

提升语音伪造检测跨语料库泛化能力,减少数据集偏差影响

Enhancing Multi-Corpus Training in SSL-Based Anti-Spoofing Models: Domain-Invariant Feature Extraction

  • 通过多任务学习与梯度反传层提取域不变特征
  • 在4个不同数据集上平均误报率降低20%
  • 适合需要跨场景鲁棒性的语音安全研究者

语音伪造检测性能在不同训练和测试语料库间常有波动。虽然多语料训练在说话人识别和语音识别中能提升鲁棒性,但我们的实验表明,该方法在伪造检测中并不总有效,甚至可能降低性能。我们推测是数据集特异性偏差影响了模型泛化能力。为此,提出域不变特征提取(IDFE)框架,采用多任务学习与梯度反传层,最小化嵌入表示中的语料特定信息。在四个不同数据集上评估,相较于基线,平均等错误率降低20%。

原文摘要 · Abstract (English)

The performance of speech spoofing detection often varies across different training and evaluation corpora. Leveraging multiple corpora typically enhances robustness and performance in fields like speaker recognition and speech recognition. However, our spoofing detection experiments show that multi-corpus training does not consistently improve performance and may even degrade it. We hypothesize that dataset-specific biases impair generalization, leading to performance instability. To address this, we propose an Invariant Domain Feature Extraction (IDFE) framework, employing multi-task learning and a gradient reversal layer to minimize corpus-specific information in learned embeddings. The IDFE framework reduces the average equal error rate by 20% compared to the baseline, assessed across four varied datasets.

语音伪造检测域不变学习自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。