arXiv:2509.21597eess.AScs.CL2025-09被引 2

构建统一音频伪造检测评估工具,覆盖31个数据集。

AUDDT: A Unified Benchmark Toolkit for Audio and Speech Deepfake Detectors

  • 整合31个音频伪造数据集,自动化评估检测器性能。
  • 在不同伪造类型和录音条件下,揭示检测器表现差异显著。
  • 适合需要验证检测器泛化能力的研究者与开发者。

随着人工智能生成内容(如音频伪造)的普及,近期大量研究聚焦于深度伪造检测技术。然而,现有基准测试仅使用有限的数据集,导致检测器在真实场景下的泛化能力难以评估。本文系统回顾了31个现有的音频伪造数据集,并提出了一个开源基准测试工具包AUDDT(https://github.com/MuSAELab/AUDDT)。该工具包旨在自动化评估预训练检测器在多种语音与非语音音频数据集上的表现,使用户能直观了解其检测器在不同伪造类型和录制条件下的优劣。我们展示了工具包的使用方法、基准构成及各类伪造子组的分布。相比现有工作,AUDDT支持跨现代欺骗方法的大规模、多样化评估,并通过全面的元数据标注实现更细致的属性级分析。基于一个广泛采用的预训练检测器,我们呈现了域内与域外检测结果,揭示了不同条件下性能存在显著波动。最后,我们分析了现有数据集的局限性及其与实际部署场景之间的差距。

原文摘要 · Abstract (English)

With the prevalence of artificial intelligence (AI)-generated content, such as audio deepfakes, a large body of recent work has focused on developing deepfake detection techniques. However, existing benchmarks employ a narrow set of datasets, leaving detector generalization to real-world conditions uncertain. In this paper, we systematically review 31 existing audio deepfake datasets and present an open-source benchmarking toolkit called AUDDT (https://github.com/MuSAELab/AUDDT). The goal of this toolkit is to automate the evaluation of pretrained detectors across a wide range of speech and non-speech audio datasets, giving users direct feedback on the advantages and shortcomings of their deepfake detectors under diverse manipulation types and recording conditions. We start by showcasing the usage of the developed toolkit, the composition of our benchmark, and the breakdown of different deepfake subgroups. Next, we highlight how AUDDT differs from existing benchmarking efforts by enabling large-scale, diverse evaluation across modern spoofing methods and richer attribute-level analysis through comprehensive metadata annotation. Using a widely adopted pretrained deepfake detector, we present in- and out-of-domain detection results, revealing notable performance variability across different conditions and audio manipulation types. Lastly, we also analyze the limitations of these existing datasets and their gaps relative to practical deployment scenarios.

音频伪造检测基准AI安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。