arXiv:2411.00121cs.SDcs.AI2024-11ICLR被引 22

提出新方法提升假音频检测模型抗干扰能力,显著增强真实场景下的识别效果。

I Can Hear You: Selective Robust Training for Deepfake Audio Detection

  • 通过高频选择性对抗训练,聚焦易被攻击的高频率特征
  • 在污染和攻击样本上准确率提升29.3%,干净样本提升7.7%
  • 构建130万样本的公开数据集DeepFakeVox-HQ,覆盖14种来源

人工智能生成语音技术的发展加剧了假音频检测的挑战,威胁到诈骗与虚假信息传播。为此,我们构建了迄今最大的公开语音数据集DeepFakeVox-HQ,包含130万条样本,其中27万条为来自14种不同来源的高质量假音频。尽管现有检测器在以往数据上表现良好,但在本数据集上性能显著下降,尤其在现实噪声和对抗攻击下。我们系统分析了提升模型鲁棒性的因素,发现多样化语音增强有益。进一步发现最优检测模型依赖人类不可感知的高频特征,易被攻击者操纵。为此,我们提出F-SAT:频率选择性对抗训练方法,专注于高频成分。实验表明,使用该数据集可使基线模型性能提升33%;采用我们的鲁棒训练后,在干净样本上准确率提高7.7%,在受损和受攻击样本上提升29.3%,优于当前最佳模型RawNet3。

原文摘要 · Abstract (English)

Recent advances in AI-generated voices have intensified the challenge of detecting deepfake audio, posing risks for scams and the spread of disinformation. To tackle this issue, we establish the largest public voice dataset to date, named DeepFakeVox-HQ, comprising 1.3 million samples, including 270,000 high-quality deepfake samples from 14 diverse sources. Despite previously reported high accuracy, existing deepfake voice detectors struggle with our diversely collected dataset, and their detection success rates drop even further under realistic corruptions and adversarial attacks. We conduct a holistic investigation into factors that enhance model robustness and show that incorporating a diversified set of voice augmentations is beneficial. Moreover, we find that the best detection models often rely on high-frequency features, which are imperceptible to humans and can be easily manipulated by an attacker. To address this, we propose the F-SAT: Frequency-Selective Adversarial Training method focusing on high-frequency components. Empirical results demonstrate that using our training dataset boosts baseline model performance (without robust training) by 33%, and our robust training further improves accuracy by 7.7% on clean samples and by 29.3% on corrupted and attacked samples, over the state-of-the-art RawNet3 model.

假音频检测对抗训练语音安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。