首个音频深度伪造检测综合基准,助力模型性能公正评估
Speech DF Arena: A Leaderboard for Speech DeepFake Detection Models
- 构建覆盖14个数据集的统一评测框架
- 多系统在跨域场景下错误率超30%,暴露泛化短板
- 适合研究者与开发者提升模型鲁棒性
随着先进语音深度伪造生成技术的发展,音频深度伪造检测也取得显著进展。然而,缺乏标准化且全面的评测基准。为此,我们提出Speech DeepFake(DF)Arena,首个面向音频深度伪造检测的综合性基准。该平台提供统一评测工具包,涵盖14个多样化数据集与攻击场景,采用标准化评估指标与流程,确保结果可复现、透明。同时设有排行榜,用于对比与排名各检测系统,助力研究人员与开发者提升模型可靠性与鲁棒性。包含12个开源及3个专有检测系统。研究发现多个系统在跨域场景下等错误率(EER)超过30%,凸显全面跨域评估的重要性。排行榜托管于Hugging Face,代码工具包可在GitHub获取。
原文摘要 · Abstract (English)
Parallel to the development of advanced deepfake audio generation, audio deepfake detection has also seen significant progress. However, a standardized and comprehensive benchmark is still missing. To address this, we introduce Speech DeepFake (DF) Arena, the first comprehensive benchmark for audio deepfake detection. Speech DF Arena provides a toolkit to uniformly evaluate detection systems, currently across 14 diverse datasets and attack scenarios, standardized evaluation metrics and protocols for reproducibility and transparency. It also includes a leaderboard to compare and rank the systems to help researchers and developers enhance their reliability and robustness. We include 14 evaluation sets, 12 state-of-the-art open-source and 3 proprietary detection systems. Our study presents many systems exhibiting high EER in out-of-domain scenarios, highlighting the need for extensive cross-domain evaluation. The leaderboard is hosted on Huggingface1 and a toolkit for reproducing results across the listed datasets is available on GitHub.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。