结合上下文和字幕可显著提升对名人音频伪造的检测效果
Context and Transcripts Improve Detection of Deepfake Audios of Public Figures
- 引入上下文与字幕信息,改进音频伪造检测模型
- 检测性能提升最高达37.58%(F1值),抗欺骗攻击能力更强
- 适合安全、媒体审核等需要高可靠性的场景
人类在判断信息真伪时会借助上下文。然而现有音频深度伪造检测器仅分析音频文件,忽略上下文与字幕信息。本文构建并分析了由超过70名记者自2024年初贡献的255个名人伪造音频数据集(JDD),并生成了已故名人合成音频数据集(SYN)。提出一种基于上下文的音频深度伪造检测新架构(CADD)。在ITW与P²V两个大规模数据集上评估显示,充分的上下文或字幕可显著提升检测效果:多个基线模型与传统分类器的F1分数提升5%-37.58%,AUC提升3.77%-42.79%,EER降低6.17%-47.83%。CADD在面对5种对抗性逃避策略时表现出更强鲁棒性,平均性能下降仅-0.71%。代码、模型与数据集可在项目页面获取(审稿期间受限)。
原文摘要 · Abstract (English)
Humans use context to assess the veracity of information. However, current audio deepfake detectors only analyze the audio file without considering either context or transcripts. We create and analyze a Journalist-provided Deepfake Dataset (JDD) of 255 public deepfakes which were primarily contributed by over 70 journalists since early 2024. We also generate a synthetic audio dataset (SYN) of dead public figures and propose a novel Context-based Audio Deepfake Detector (CADD) architecture. In addition, we evaluate performance on two large-scale datasets: ITW and P$^2$V. We show that sufficient context and/or the transcript can significantly improve the efficacy of audio deepfake detectors. Performance (measured via F1 score, AUC, and EER) of multiple baseline audio deepfake detectors and traditional classifiers can be improved by 5%-37.58% in F1-score, 3.77%-42.79% in AUC, and 6.17%-47.83% in EER. We additionally show that CADD, via its use of context and/or transcripts, is more robust to 5 adversarial evasion strategies, limiting performance degradation to an average of just -0.71% across all experiments. Code, models, and datasets are available at our project page: https://sites.northwestern.edu/nsail/cadd-context-based-audio-deepfake-detection (access restricted during review).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。