仅用数据集身份信息提升语音伪造检测泛化能力
Leveraging Gradient Reversal Loss and Multitask Learning for Datasets-Aware Audio Deepfake Detection

- 利用数据集身份做多任务学习与梯度反转,抑制特定数据集特征
- 平均错误率降低13.14%,合并错误率降低5.32%
- 适合部署于多样且未知真实场景的语音伪造检测系统
语音合成与语音转换技术的进展对安全与隐私构成威胁,亟需可靠的深度伪造检测技术。现有系统在单一数据集上表现良好,但在跨数据集泛化上效果不佳。此前方法依赖语言、编码类型等辅助信息,但这些信息往往不全或难以获取。本文提出一种实用的、面向数据集的检测框架,仅使用天然可得的数据集身份作为监督信号,结合多任务学习与梯度反转层,实现对数据集特性的感知监督和对抗性抑制。依据2025年语音深度伪造竞技场基准协议,实验在多个评估数据集上进行,以等错误率(EER)衡量整体性能,包括平均EER与合并EER。相比基线,多任务学习使平均EER相对降低13.14%,梯度反转层使合并EER相对降低5.32%。结果表明,该方法能有效提升在异构数据集上的综合检测性能,为部署于多样且未见的真实数据提供了实用方案。
原文摘要 · Abstract (English)
Recent advances in speech synthesis and voice conversion, which pose threats to security and privacy, have underscored the need for deepfake detection technology. Although existing detection systems achieve strong performance on individual datasets, they often fail to generalize across diverse datasets. Prior methods for improving generalization, including data augmentation, adversarial training on auxiliary factors such as language or codec types, and Mixture-of-Experts (MoE), are limited by predefined augmentation coverage, difficulties in obtaining auxiliary factors, and substantial model complexity. In this work, we propose a practical dataset-aware framework for deepfake detection. Our method targets heterogeneous datasets for which auxiliary annotations such as language, codec, or spoofing method may not be consistently available. We therefore rely only on dataset identity as a naturally available supervisory signal for multitask (MT) and gradient reversal layer (GRL) training, allowing the model to investigate both dataset-aware multitask supervision and adversarial suppression of dataset-specific information. We conduct experiments following the 2025 Speech DeepFake Arena benchmark protocol, evaluating our model across multiple evaluation datasets and reporting aggregate performance in terms of Equal Error Rate (EER), including Average EER and Pooled EER. Compared with the baseline, MT reduces Average EER by 13.14% relatively, while GRL reduces Pooled EER by 5.32% relatively. These results demonstrate that our method can improve aggregate detection performance across heterogeneous evaluation datasets, offering a practical solution for deploying reliable deepfake detection systems on diverse and unseen real-world data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。