arXiv:2509.09204cs.SDcs.AI2025-09被引 5

提出新评测方法,让语音伪造检测更真实可靠。

Bona fide Cross Testing Reveals Weak Spot in Audio Deepfake Detection Systems

  • 用多样真实语音数据做交叉测试,避免样本偏倚。
  • 在9种语音类型上测试超150个合成器,性能更均衡。
  • 适合想提升检测系统真实场景鲁棒性的研究者。

语音深度伪造检测(ADD)模型通常使用混合多种合成器的数据集进行评估,性能以单一等错误率(EER)报告。然而,该方法对样本量大的合成器过度加权,导致其他合成器被低估,降低整体EER可靠性。此外,多数ADD数据集缺乏真实语音多样性,常仅包含单一环境和语音风格(如清晰朗读),难以模拟真实场景。为此,我们提出真实语音交叉测试(bona fide cross-testing)框架,引入多样化真实语音数据集,并对不同类别分别计算并聚合EER,实现更均衡的评估。我们在九类真实语音上对超过150个合成器进行了基准测试,并发布了新数据集以促进后续研究,地址:https://github.com/cyaaronk/audio_deepfake_eval。

原文摘要 · Abstract (English)

Audio deepfake detection (ADD) models are commonly evaluated using datasets that combine multiple synthesizers, with performance reported as a single Equal Error Rate (EER). However, this approach disproportionately weights synthesizers with more samples, underrepresenting others and reducing the overall reliability of EER. Additionally, most ADD datasets lack diversity in bona fide speech, often featuring a single environment and speech style (e.g., clean read speech), limiting their ability to simulate real-world conditions. To address these challenges, we propose bona fide cross-testing, a novel evaluation framework that incorporates diverse bona fide datasets and aggregates EERs for more balanced assessments. Our approach improves robustness and interpretability compared to traditional evaluation methods. We benchmark over 150 synthesizers across nine bona fide speech types and release a new dataset to facilitate further research at https://github.com/cyaaronk/audio_deepfake_eval.

语音伪造检测评估数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。