arXiv:2506.09606eess.AS2025-06被引 15

用真实世界音频伪造数据提升检测模型性能

Unmasking real-world audio deepfakes: A data-centric approach

  • 聚焦数据质量,通过清洗、裁剪和增强优化数据集
  • 在真实场景数据上将误报率降低63%
  • 适合关注实际应用的AI安全研究者

真实世界中的音频深度伪造日益增多,但现有检测系统多基于科研用途数据集评估,难以应对现实挑战。为此,本文构建了一个新型真实世界音频深度伪造数据集。分析表明,这些真实样本对最先进检测模型也构成严峻考验。不同于增加模型复杂度,本文采用数据中心范式,通过数据集清洗、修剪和增强等策略提升模型鲁棒性与泛化能力。实验结果显示,在In-the-Wild数据集上,错误等价率(EER)相对降低55%,达到1.7%;在新提出的AI4T数据集上,EER降低63%。结果凸显数据驱动方法在真实场景深度伪造检测中的巨大潜力。代码与数据见:https://github.com/davidcombei/AI4T。

原文摘要 · Abstract (English)

The growing prevalence of real-world deepfakes presents a critical challenge for existing detection systems, which are often evaluated on datasets collected just for scientific purposes. To address this gap, we introduce a novel dataset of real-world audio deepfakes. Our analysis reveals that these real-world examples pose significant challenges, even for the most performant detection models. Rather than increasing model complexity or exhaustively search for a better alternative, in this work we focus on a data-centric paradigm, employing strategies like dataset curation, pruning, and augmentation to improve model robustness and generalization. Through these methods, we achieve a 55% relative reduction in EER on the In-the-Wild dataset, reaching an absolute EER of 1.7%, and a 63% reduction on our newly proposed real-world deepfakes dataset, AI4T. These results highlight the transformative potential of data-centric approaches in enhancing deepfake detection for real-world applications. Code and data available at: https://github.com/davidcombei/AI4T.

深度伪造音频检测数据驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。